Daejeon, July 27–31
This study uses the predictive mechanism of autoregressive language models to simulate how repeated exposure to a dominant discourse might reshape historically specific textual expectations. The adopted approach is inspired by research pointing to alignments between semantic representations in large language models (LLMs) and human brain activity, suggesting that LLMs can offer insights into human narrative comprehension (Caucheteux and King 2022; Kumar et al. 2024; Tikochinski 2025) and historical behavioral science (Varnum et al. 2024). While acknowledging the important differences between artificial and biological cognition, I leverage LLMs to revisit a century-old dialogue between probabilists and humanists, a conversation initiated by Andrey Markov’s analysis of Pushkin’s Eugene Onegin and later reframed by Claude Shannon, who cast humans as predictive language models implicitly aware of language statistics (Markov 2006; Shannon 1951). The present study uses perplexity, a measure of a model’s predictive certainty derived from cross-entropy loss (Jelinek et al. 1977), to quantify a model’s “surprise” when encountering a given text. A low perplexity score indicates high predictability, whereas a high score signals unexpectedness.
While stylometrists traditionally favored tangible features like function words and syntactic structures (Lutosławski 1898; Mosteller and Wallace 1963; Burrows 1987; Herrmann, Dalen-Oskam, and Schöch 2015; Rybicki, Eder, and Hoover 2016; Evert et al. 2017), the advent of LLMs has brought predictability to the forefront. Perplexity emerged as a key marker for differentiating machine from human text (Holtzman et al. 2020, Mitchell et al. 2023), though the increasing sophistication of LLMs has prompted a shift towards other methods like watermarking for provenance verification (Kirchenbauer et al. 2023; Zhao et al. 2023). Concurrently, studies in the computational humanities have begun to explore perplexity’s stylometric potential, using it to distinguish canonical from non-canonical novels (Wu et al. 2024), machine from human translation (Bizzoni et al. 2020), and political from novelistic prose (Kurzynski 2023, 2024), signaling a broader turn towards cognitive and informational processing in digital literary analysis.
Perplexity as a metric offers new insights into the tension between the familiar and the unexpected, a dynamic central to Viktor Shklovsky’s theory of defamiliarization (ostranenie). Shklovsky argued that art’s function is to combat the “automatization of perception,” a process whereby habit renders life routine and unfelt. Through defamiliarization, art “increases the complexity and duration of perception” and “recovers the sensation of life” by presenting familiar forms in unfamiliar ways (Shklovsky 2017). This framework finds empirical support in modern cognitive science, from eye-tracking experiments on unexpected words (McDonald and Shillock 2004) to fMRI scans revealing heightened brain activity when processing unusual sentence structures, consistent with theories of the brain as a predictive organ (Clark 2013; Bohrn et al. 2012). Style can thus be understood not only as a static property of the text but as a source of dynamic engagement with a reader’s predictive faculties. Accordingly, I use “cognitive freedom” not as a psychological essence but as a formal property of aesthetic experience: the capacity of language to keep perception open by resisting complete predictability.
The core of this project is a two-stage simulation designed to model, in simplified computational form, how repeated exposure to dominant political language reshaped textual expectations in post-1949 China (Link 2013). First, a GPT2-style Transformer model (223M parameters, 16 layers, 16 attention heads) was pre-trained from scratch on a large, general-purpose corpus of modern Chinese texts (the “4-5” split of FineWeb Edu Chinese V2.1, containing ~70 GB of content filtered for educational value) (Yu et al. 2025). This base model serves as a general-language baseline with broad statistical exposure to modern Chinese. A custom character-level tokenizer was constructed, treating each Chinese character as a distinct token for a more granular measurement of surprise. The training corpus was screened for potential data leakage from test corpora (described below) using MD5 hashing and literal string matching on unique 13-grams.
In the second stage, the base model was fine-tuned exclusively on the Selected Works of Mao Zedong 毛泽东选集 for five epochs. This repeated training simulates intense exposure to a single, dominant idiolect, creating a specialized reader of “Maospeak”—the militant, ideologically charged language style that saturated Chinese public life during the Mao era (1949-1976) (Li 1998; Leese 2011). While Mao himself was a creative rhetorician, the institutionalization of Maoism transformed his words into instruments of mass indoctrination and ritualistic recitation, particularly during the Cultural Revolution, when proficiency in Maospeak could become a condition of social, political, and sometimes physical survival (Ji 2004; Schoenhals 2007).
Figure 1. Per‑character perplexity plot for an excerpt from the Selected Works of Mao Zedong.
By tracking the decrease in the model’s perplexity on the Mao corpus, the study identifies the core phraseology that becomes “automatized” in the Shklovskian sense. The phrases with the most significant drop in perplexity are central to the era’s political machinery, including canonical lists of class adversaries (“landlords, rich peasants, counter-revolutionaries, bad elements, and Rightists” 地、富、反、坏、右), labels for political targets (“unrepentant capitalist-roader” 不肯改悔的走资派), and formulaic rhetoric (“resolutely, thoroughly, wholly, and completely annihilate” 坚决、彻底、干净、全部地消灭掉). Visualizing these as “perplexity landscapes” reveals a characteristic pattern: the first character of a key phrase often generates a high perplexity spike, but once revealed, it constrains subsequent possibilities, causing perplexity to cascade into deep, low-perplexity “canyons” of predictable slogans (Figure 1). While this phenomenon is a general feature of any language, what is distinctive about Maospeak is the length and frequency of such rigid sequences. The formulaic nature of Maospeak was further quantified through an n-gram entropy analysis, which confirmed that the Mao corpus is significantly more repetitive and less varied (i.e., lower in Shannon entropy) than a corpus of modern Chinese novels at every n-gram size tested, with the most pronounced difference observed at lower n-grams (2-6), which contain the bulk of political vocabulary.
This process offers a computational analogue for Shklovskian familiarization, where a powerful discourse creates a predictable linguistic universe, fostering a “psychological atmosphere of control, certainty, and patent purpose” (Tsur 2008). However, such enforced familiarization comes at a cognitive cost. The average perplexity of the model on a corpus of 100 modern Chinese novels (20世纪中文小说100强) increased with each epoch of fine-tuning on Mao's works. I describe this effect as “cognitive overfitting”: specialization in one discourse reduces the model’s tolerance for stylistic alternatives. This trade-off highlights the opposing principles at play, as an analysis of excerpts from three major Chinese novelists illustrates how literature operates through defamiliarization.
In Zhang Wei’s The Ancient Ship 古船 (1987), for instance, the model’s perplexity drops significantly on embedded Maoist-era slogans such as “Those who are not afraid of being cut to a thousand pieces dare to pull the emperor off his horse” 舍得一身剐, 敢把皇帝拉下马, even as the surrounding narrative remains surprising to the fine-tuned model. This demonstrates the method’s ability to isolate the intertextual presence of a familiar, dominant discourse. Conversely, in Lilian Lee’s Farewell My Concubine 霸王别姬 (1985), the analysis highlights how literary language creates surprise. The highest perplexity spikes occur on creative juxtapositions like “a ferocious yawn” (凶狠地打哈欠) and on pre-modern phrasings. As the model fine-tunes on the functional register of Maospeak, its perplexity on these literary and classical expressions increases, showing how its linguistic worldview has narrowed. Finally, Dung Kai-cheung’s postmodern Hong Kong novel Works and Creations 天工开物・栩栩如真 (2005) exemplifies defamiliarization through linguistic disruption. The author deliberately inserts Cantonese vernacular characters (e.g., 嗰, 系, 既) into standard written Chinese (Snow 2004), which generates sharp perplexity spikes, disrupting the automatized perception of the standard language and forcing a confrontation with the text’s cultural specificity (Figure 2).
Figure 2. A per‑character perplexity plot for an excerpt from Dung Kai‑cheung’s Works and Creations.
Ultimately, this study suggests that style can be viewed as a “cognitive signature”: a text’s unique strategy for managing a reader’s attention by orchestrating the tension between predictability and surprise. The perplexity arc of a text, mapping these oscillations between the expected and the unexpected, can be seen as a cognitive counterpart to the “emotional arcs” identified in sentiment analysis (Reagan et al. 2016; Elkins 2022). While engineered political language like Maospeak seeks to minimize perplexity and reinforce ideology through low-entropy patterns, literary language thrives on generating “non-anomalous surprise” (Hogan 2016). Literature uses a predictable backdrop of narrative and linguistic convention to make its high-perplexity focal points (e.g., a startling metaphor, a disruptive dialect, a novelistic event) more impactful. As such, cognitive stylometry allows us to move beyond models of style as a detachable surface feature, towards a more integrated, cognitive-formalist understanding where form and content are inseparable. Beyond sinology, this study also speaks to a broader contemporary problem: the mechanization and homogenization of language under ideological and algorithmic regimes alike, as well as the importance of linguistic diversity for resisting the narrowing of perception, expression, and imagination.