ArXiv paper proposes five-part rubric marketers can use to judge AI copy
Learn how the 'Literary Axioms' framework helps marketers evaluate AI-generated creative using five key dimensions for better content quality.
Hannah Vogel ·
In an arXiv preprint updated as version 2609.25504v2, the authors propose a conceptual framework they call “Literary Axioms” that ties a work’s creation, interpretation and evaluation together, and offer five dimensions—coherence, contextual reach, distinctiveness, realization fidelity, and interpretive richness—for judging texts. This is, so far, single-source — arXiv only, with no independent confirmation and no one on the record. The preprint includes a concrete implementation comprising 1,455 Chinese-language records, 1,464 mappings to 149 works, and 472 typed relationships, as well as a descriptive audit that notes gaps in reuse, context coding and provenance. The authors state their predictions await reader and creative writing studies. According to the preprint, the framework’s unit of analysis is a “literary axiom”: a provisional organizing premise whose scope and textual realization can be examined, selected and combined by writers, reconstructed by readers, and evaluated on the five dimensions named above. arXiv
The framework formalizes how a text’s premise is selected, realized and reconstructed
The preprint distinguishes four stages: an open axiom space (the set of possible organizing premises), a selected configuration (the subset a writer chooses), its relational organization (how those premises interact across characters, events, language and form), and situated reader reconstructions (the plurality of grounded ways readers rebuild meaning from the text). From this, it defines coherence as organized compatibility or conflict among premises; originality as a change in claims, relations or realization; and interpretive richness as a plurality of textually grounded reconstructions. The five proposed dimensions for evaluation—coherence, contextual reach, distinctiveness, realization fidelity, and interpretive richness—derive from this structure. The implementation’s descriptive audit aims to establish structural integrity while acknowledging gaps in reuse, context coding, and provenance, and the authors explicitly present five empirical predictions with observations that would count against them. In other words, this is a theory-led framework with a falsifiability posture, but not yet a validated measurement system. arXiv
Why this matters to marketing now: it offers measurable acceptance criteria for AI outputs
Marketers increasingly procure or build systems that generate copy, scripts and campaign concepts; acceptance criteria for these outputs are often limited to surface checks (style guides, banned words, readability scores). The preprint’s five dimensions map onto practical review questions: coherence (does the piece maintain an internally consistent premise across headlines, body and CTA), contextual reach (does it accommodate the audience segments and channels specified in the brief), distinctiveness (does it demonstrate a clear differentiation claim or tone), realization fidelity (does the language, imagery and structure execute the chosen premises faithfully), and interpretive richness (does the piece support multiple, grounded readings without collapsing into vagueness). Because the framework ties these to an articulated configuration of “premises,” it gives CMOs and creative leads a common vocabulary to embed in RFPs, QA checklists and vendor scorecards for generative tools, beyond generic “on-brand” judgments. For procurement, that means clearer acceptance tests in statements of work, and for sales and customer success on the vendor side, a clearer articulation of how a system will be evaluated in pilot and renewal. arXiv
The implementation is inspectable, but the evidence is early and scoped
The preprint’s authors describe a concrete artifact with 1,455 Chinese-language records, 1,464 mappings to 149 works, and 472 typed relationships, and they provide a descriptive audit noting structural integrity as well as gaps in reuse, context coding, and provenance. They also state that their predictions await reader and creative writing studies. For enterprise operators, that combination matters: the dataset’s language and scope signal that immediate generalization to English-language ad copy or global brand creative is unproven, and the authors’ own audit foregrounds limitations in metadata reuse and context that would need resolution before automation is built on top. Still, the presence of typed relationships and a relational organization suggests the framework can be represented in data structures marketers already use (taxonomy tables, content models, CMS schemas), making it feasible to trial as a governance layer even while empirical validation proceeds. arXiv
Governance and audit: a conceptual spine for AI creative that can be inspected later
Internal AI governance programs require explainability and auditability of outputs, but most controls focus on inputs (training data, prompts) and surface outcomes (prohibited content). By proposing that a text’s “premises” be explicitly selected and that their relational organization be documented, the preprint supplies a way to generate an audit log specific to creative work: which premises were declared in the brief, how they relate, where in the text they are realized, and how reviewers reconstructed them. The five dimensions then serve as checkpoints that can be attested to in QA or post-mortems. Even if no regulator mandates such documentation, large buyers already ask for model cards, content risk attestations and explainability statements; extending that practice to the creative output layer could shorten procurement cycles and clarify liability handoffs between vendor and buyer. If adopted, this moves some risk from subjective taste into documented criteria, which changes how agencies pitch and how software vendors sell into marketing: success becomes not just “did it perform,” but “did it adhere to declared premises in ways we can inspect.” arXiv
The skeptic’s read: scoring art invites spurious precision and cultural bias
There is an obvious counter: creative evaluation resists checklists, and the risk of overfitting a rubric to heterogeneous audiences is real. The preprint itself underscores that interpretive richness matters, which implies variability across readers and contexts; locking a single configuration into an enterprise rubric could suppress novelty and lead to safe sameness. The dataset’s Chinese-language focus also raises portability questions: claims about contextual reach or distinctiveness are culturally embedded, and a rubric derived from one corpus may not transpose cleanly to another. Finally, while the framework is falsifiable on paper, the preprint states its predictions await reader and creative writing studies; until those arrive, operationalizing the five dimensions risks giving a false sense of rigor. That skeptical stance doesn’t negate the utility of a shared vocabulary—but it does argue for treating the five dimensions as conversation aids and acceptance scaffolds, not as a scoring regime whose numbers drive compensation or vendor penalties yet. arXiv
What changes first: RFP language, QA forms and vendor demos
If this framework is to matter commercially, the earliest observable shifts will be in documents, not dashboards. RFPs that today ask vendors to match brand voice may begin to request how proposed systems will support “premise selection” and provide controls for coherence and realization fidelity; agency statements of work may add deliverables for premise documentation and reader reconstruction logs. On the vendor side, demos may add features to tag and trace premises through content variants, then report on the five dimensions as optional fields in CMS or DAM integrations. Internal creative QA forms could be revised to include interpretive richness as a deliberate checkpoint, making teams explicit about when plurality is desired and where it should be constrained. If these appear, they will show up in procurement timelines and renewal criteria, subtly shifting how marketing buys AI software and how vendors train their sales engineers to argue capability and accountability. arXiv
What to watch over the next two quarters
Three signals will tell whether this research is more than a theoretical curiosity for enterprise marketing. First, whether major brand or agency RFP templates publicly add language that mirrors the five dimensions, or ask vendors to document premise selection and relational organization of outputs. Second, whether enterprise content platforms and QA tools add fields or reports labeled along these dimensions or link to the arXiv framework in whitepapers or documentation. Third, whether any reader or creative writing study tests the framework’s predictions and reports outcomes that validate or challenge its use as an evaluation scaffold; a null result would argue for caution in productizing it, while supportive findings would harden buyers’ expectations that vendors measure what they claim to support. In each case, the paper’s own audit—and its candid statement that predictions await study—sets the bar: marketing operators should expect inspectable artifacts, clearly stated limits, and a path to falsification before they commit the framework to performance management. arXiv