Daejeon, July 27–31
This study examines the interpretive behavior of LLMs as literary critics by profiling their critical practice. Recent scholarship has explored LLMs' potential in critical domains, demonstrating their capacity for structuralist analysis (Dong et al. 2025), art critique (Arita et al. 2025), creative interpretive assemblages (O'Halloran 2024), speculative posthumanist writing (Bardzell / Ghajargar 2026), adversarial critical reading (Floridi et al. 2025), probing literary understanding (Jannidis et al. 2025), and argumentative interpretation evaluation (Pichler et al. 2026). The structural and semantic dynamics underlying algorithmic interpretation itself, however, remain under-explored. This research analyzes the algorithmic characteristics of LLM interpretation, examining patterns of critical lens activation, evidence selection, and semantic distribution under varied prompt conditions. Through lens-frequency analysis, citation analysis, and vector semantics, the study aims to map the characteristics and boundaries of algorithmic criticism. Beyond performance metrics, this study addresses the humanities question of algorithmic influence on interpretive diversity and scholarly knowledge production.
We focus on three canonical English-language poems chosen for their structural complexity and divergent histories of critical reception: T. S. Eliot's The Waste Land, Christina Rossetti's Goblin Market, and Allen Ginsberg's Howl. The research is organized around three guiding questions to systematically chart the interpretive operations of LLMs as literary critics. First, which critical lenses does an LLM preferentially activate? Repeated mobilization of specific lenses would suggest that LLMs do not merely reproduce existing critical resources but operate as agents reshaping the distribution of critical perspectives at the field level. Second, how do varied prompt conditions reshape these activation patterns? Although LLMs internalize a broad repertoire of critical lenses, their generative mechanisms may yield distinctive deployment patterns under different prompt cues. Third, how does the activation profile shift across analytical targets? Lens activation may be a target-independent regularity or a context-sensitive phenomenon shaped by genre, period, and thematic properties of the source text.
The methodology follows a sequential pipeline: data construction, AI-critique generation, and quantitative analysis. For the human baseline, we constructed a corpus of 120 articles centered on The Waste Land by querying the poem's title in dedicated venues: T. S. Eliot Studies Annual (2017–), The Journal of the TS Eliot Society (2024–2025), and Yeats-Eliot Review (1974–2015). This Waste Land corpus serves as the text-specific human reference for calibrating AI critical practice, while Goblin Market and Howl, drawn from different periods, genres, and thematic registers, extend the AI corpus to additional analytical targets that anchor the cross-target analysis. Using Gemini 3.1 Pro (high-reasoning mode), we extracted specific lines cited and critical lenses applied in each article to establish a baseline of scholarly patterns.
To generate the AI corpus, we utilize two state-of-the-art reasoning models, GPT-5.2 and Claude Opus 4.5 (both in high-reasoning mode), prompting each to produce literary criticism on the three poems. To surface the breadth of lenses LLMs mobilize, we employ a five-prompt strategy: a baseline condition and four lexical-injection conditions populated by adjectives from the novelty semantic field: fresh, novel, innovative, groundbreaking. Loosely ordered in intensity, these modifiers serve as a methodological device for surfacing latent interpretive frames while pushing the models toward the interpretive originality conventionally taken to be a central aim of literary scholarship. For each model–prompt combination, we generate 40 critical essays per poem, yielding 1,200 essays in total.
In the analysis stage, we employ Gemini 3.1 Pro (high-reasoning mode)—a model family distinct from those used in generation—to extract cited lines and critical lenses from each essay, mitigating potential self-extraction bias. The analysis proceeds along three complementary tracks. First, for The Waste Land, we assess the LLM baseline's line-level citation distribution against the human reference corpus using Jensen–Shannon divergence and permutation testing. Second, across all three poems and all five prompt conditions, we apply the same procedure to compare lens-frequency distributions across the AI corpus and against the Waste Land human reference. Third, we encode each essay into 2048-dimensional vectors using the voyage-4-large embedding model and apply Principal Component Analysis (PCA) to visualize the topological scope of the resulting interpretive landscapes, comparing the dispersion of semantic vectors across models, prompt conditions, and the human reference corpus using mean pairwise cosine distance. To validate these patterns, we conduct close readings of representative essays and outliers.
Findings indicate distinctive structural patterns in algorithmic criticism. First, AI models exhibit a higher rate of textual citation than human critics, and their citation patterns are markedly more linear, closely following the poem’s line order. This linearity suggests that algorithmic interpretation tracks the source text's surface progression, whereas human criticism reorganizes textual evidence around thematic and argumentative clusters. Second, AI baseline criticism mirrors the most prevalent theoretical orientations in the human corpus: dominant lenses such as theological criticism, formalism, and intertextuality persist at non-trivial proportions even under novelty-injected prompts, indicating transfer from the critical literature embedded in training data. Third, when novelty modifiers are introduced, LLMs generate their own new dominant lenses rather than drawing more widely from established critical traditions, a pattern we term theory homogenization. The models converge on theoretical attractors distinct from the human distribution, each with a distinctive discursive signature. Although not strictly consistent, increasing the intensity of the novelty adjectives (e.g., from fresh to groundbreaking) appears to amplify the prominence of these emergent lenses, a tendency we describe as theory amplification. PCA-based semantic analysis corroborates these patterns: human–AI distances exceed inter-model distances in embedding space, each model retains a distinct semantic signature per poem, and inter-condition distances grow with novelty intensity, consistent with theory amplification.
This study characterizes the distinctive epistemic profile of LLMs as literary critics, aiming to increase understanding of their interpretive behavior and to support more informed choices about their use in research. Whereas the critical framework surrounding a given literary work has historically coalesced through the cumulative, gradual validation and contestation by a scholarly community, the LLM-generated framework emerges instantaneously and gravitates toward homogenization, even when explicitly prompted toward novelty. We further suggest that the observed theory homogenization may not merely be a byproduct of training data but a reflection of the model's architectural tendency toward probabilistic convergence, a structural dynamic we propose to examine in this paper.