DH 2026

Daejeon, July 27–31

Wed, July 2909:00–10:30S014105
Long Paper

Perplexity as a Measure of Cognitive Freedom: An LLM-Based Study of Defamiliarization in Modern Chinese

Maciej Kurzynski
Lingnan University, Hong Kong S.A.R. · maciej.kurzynski@ln.edu.hk

This study uses the predictive mechanism of autoregressive language models to simulate how repeated exposure to a dominant discourse might reshape historically specific textual expectations. The adopted approach is inspired by research pointing to alignments between semantic representations in large language models (LLMs) and human brain activity, suggesting that LLMs can offer insights into human narrative comprehension (Caucheteux and King 2022; Kumar et al. 2024; Tikochinski 2025) and historical behavioral science (Varnum et al. 2024). While acknowledging the important differences between artificial and biological cognition, I leverage LLMs to revisit a century-old dialogue between probabilists and humanists, a conversation initiated by Andrey Markov’s analysis of Pushkin’s Eugene Onegin and later reframed by Claude Shannon, who cast humans as predictive language models implicitly aware of language statistics (Markov 2006; Shannon 1951). The present study uses perplexity, a measure of a model’s predictive certainty derived from cross-entropy loss (Jelinek et al. 1977), to quantify a model’s “surprise” when encountering a given text. A low perplexity score indicates high predictability, whereas a high score signals unexpectedness.

While stylometrists traditionally favored tangible features like function words and syntactic structures (Lutosławski 1898; Mosteller and Wallace 1963; Burrows 1987; Herrmann, Dalen-Oskam, and Schöch 2015; Rybicki, Eder, and Hoover 2016; Evert et al. 2017), the advent of LLMs has brought predictability to the forefront. Perplexity emerged as a key marker for differentiating machine from human text (Holtzman et al. 2020, Mitchell et al. 2023), though the increasing sophistication of LLMs has prompted a shift towards other methods like watermarking for provenance verification (Kirchenbauer et al. 2023; Zhao et al. 2023). Concurrently, studies in the computational humanities have begun to explore perplexity’s stylometric potential, using it to distinguish canonical from non-canonical novels (Wu et al. 2024), machine from human translation (Bizzoni et al. 2020), and political from novelistic prose (Kurzynski 2023, 2024), signaling a broader turn towards cognitive and informational processing in digital literary analysis.

Perplexity as a metric offers new insights into the tension between the familiar and the unexpected, a dynamic central to Viktor Shklovsky’s theory of defamiliarization (ostranenie). Shklovsky argued that art’s function is to combat the “automatization of perception,” a process whereby habit renders life routine and unfelt. Through defamiliarization, art “increases the complexity and duration of perception” and “recovers the sensation of life” by presenting familiar forms in unfamiliar ways (Shklovsky 2017). This framework finds empirical support in modern cognitive science, from eye-tracking experiments on unexpected words (McDonald and Shillock 2004) to fMRI scans revealing heightened brain activity when processing unusual sentence structures, consistent with theories of the brain as a predictive organ (Clark 2013; Bohrn et al. 2012). Style can thus be understood not only as a static property of the text but as a source of dynamic engagement with a reader’s predictive faculties. Accordingly, I use “cognitive freedom” not as a psychological essence but as a formal property of aesthetic experience: the capacity of language to keep perception open by resisting complete predictability.

The core of this project is a two-stage simulation designed to model, in simplified computational form, how repeated exposure to dominant political language reshaped textual expectations in post-1949 China (Link 2013). First, a GPT2-style Transformer model (223M parameters, 16 layers, 16 attention heads) was pre-trained from scratch on a large, general-purpose corpus of modern Chinese texts (the “4-5” split of FineWeb Edu Chinese V2.1, containing ~70 GB of content filtered for educational value) (Yu et al. 2025). This base model serves as a general-language baseline with broad statistical exposure to modern Chinese. A custom character-level tokenizer was constructed, treating each Chinese character as a distinct token for a more granular measurement of surprise. The training corpus was screened for potential data leakage from test corpora (described below) using MD5 hashing and literal string matching on unique 13-grams.

In the second stage, the base model was fine-tuned exclusively on the Selected Works of Mao Zedong 毛泽东选集 for five epochs. This repeated training simulates intense exposure to a single, dominant idiolect, creating a specialized reader of “Maospeak”—the militant, ideologically charged language style that saturated Chinese public life during the Mao era (1949-1976) (Li 1998; Leese 2011). While Mao himself was a creative rhetorician, the institutionalization of Maoism transformed his words into instruments of mass indoctrination and ritualistic recitation, particularly during the Cultural Revolution, when proficiency in Maospeak could become a condition of social, political, and sometimes physical survival (Ji 2004; Schoenhals 2007).

Figure 1. Per‑character perplexity plot for an excerpt from the Selected Works of Mao Zedong.

By tracking the decrease in the model’s perplexity on the Mao corpus, the study identifies the core phraseology that becomes “automatized” in the Shklovskian sense. The phrases with the most significant drop in perplexity are central to the era’s political machinery, including canonical lists of class adversaries (“landlords, rich peasants, counter-revolutionaries, bad elements, and Rightists” 地、富、反、坏、右), labels for political targets (“unrepentant capitalist-roader” 不肯改悔的走资派), and formulaic rhetoric (“resolutely, thoroughly, wholly, and completely annihilate” 坚决、彻底、干净、全部地消灭掉). Visualizing these as “perplexity landscapes” reveals a characteristic pattern: the first character of a key phrase often generates a high perplexity spike, but once revealed, it constrains subsequent possibilities, causing perplexity to cascade into deep, low-perplexity “canyons” of predictable slogans (Figure 1). While this phenomenon is a general feature of any language, what is distinctive about Maospeak is the length and frequency of such rigid sequences. The formulaic nature of Maospeak was further quantified through an n-gram entropy analysis, which confirmed that the Mao corpus is significantly more repetitive and less varied (i.e., lower in Shannon entropy) than a corpus of modern Chinese novels at every n-gram size tested, with the most pronounced difference observed at lower n-grams (2-6), which contain the bulk of political vocabulary.

This process offers a computational analogue for Shklovskian familiarization, where a powerful discourse creates a predictable linguistic universe, fostering a “psychological atmosphere of control, certainty, and patent purpose” (Tsur 2008). However, such enforced familiarization comes at a cognitive cost. The average perplexity of the model on a corpus of 100 modern Chinese novels (20世纪中文小说100强) increased with each epoch of fine-tuning on Mao's works. I describe this effect as “cognitive overfitting”: specialization in one discourse reduces the model’s tolerance for stylistic alternatives. This trade-off highlights the opposing principles at play, as an analysis of excerpts from three major Chinese novelists illustrates how literature operates through defamiliarization.

In Zhang Wei’s The Ancient Ship 古船 (1987), for instance, the model’s perplexity drops significantly on embedded Maoist-era slogans such as “Those who are not afraid of being cut to a thousand pieces dare to pull the emperor off his horse” 舍得一身剐, 敢把皇帝拉下马, even as the surrounding narrative remains surprising to the fine-tuned model. This demonstrates the method’s ability to isolate the intertextual presence of a familiar, dominant discourse. Conversely, in Lilian Lee’s Farewell My Concubine 霸王别姬 (1985), the analysis highlights how literary language creates surprise. The highest perplexity spikes occur on creative juxtapositions like “a ferocious yawn” (凶狠地打哈欠) and on pre-modern phrasings. As the model fine-tunes on the functional register of Maospeak, its perplexity on these literary and classical expressions increases, showing how its linguistic worldview has narrowed. Finally, Dung Kai-cheung’s postmodern Hong Kong novel Works and Creations 天工开物・栩栩如真 (2005) exemplifies defamiliarization through linguistic disruption. The author deliberately inserts Cantonese vernacular characters (e.g., 嗰, 系, 既) into standard written Chinese (Snow 2004), which generates sharp perplexity spikes, disrupting the automatized perception of the standard language and forcing a confrontation with the text’s cultural specificity (Figure 2).

Figure 2. A per‑character perplexity plot for an excerpt from Dung Kai‑cheung’s Works and Creations.

Ultimately, this study suggests that style can be viewed as a “cognitive signature”: a text’s unique strategy for managing a reader’s attention by orchestrating the tension between predictability and surprise. The perplexity arc of a text, mapping these oscillations between the expected and the unexpected, can be seen as a cognitive counterpart to the “emotional arcs” identified in sentiment analysis (Reagan et al. 2016; Elkins 2022). While engineered political language like Maospeak seeks to minimize perplexity and reinforce ideology through low-entropy patterns, literary language thrives on generating “non-anomalous surprise” (Hogan 2016). Literature uses a predictable backdrop of narrative and linguistic convention to make its high-perplexity focal points (e.g., a startling metaphor, a disruptive dialect, a novelistic event) more impactful. As such, cognitive stylometry allows us to move beyond models of style as a detachable surface feature, towards a more integrated, cognitive-formalist understanding where form and content are inseparable. Beyond sinology, this study also speaks to a broader contemporary problem: the mechanization and homogenization of language under ideological and algorithmic regimes alike, as well as the importance of linguistic diversity for resisting the narrowing of perception, expression, and imagination.

References
  1. Bizzoni, Yuri, Juzek, Tom S., España-Bonet, Cristina, et al. 2020. How Human is Machine Translationese? Comparing Human and Machine Translations of Text and Speech. In Proceedings of the 17th International Conference on Spoken Language Translation, 280–290. Association for Computational Linguistics. https://aclanthology.org/2020.iwslt-1.34/.
  2. Bohrn, Urs V., Ulrike Altmann, Oliver Lubrich, et al. 2012. “Old Proverbs in New Skins: An fMRI Study on Defamiliarization.” Frontiers in Psychology 3:204.
  3. Burrows, John F. 1987. “Word-Patterns and Story-Shapes: The Statistical Analysis of Narrative Style.” Literary and Linguistic Computing 2 (2): 61– 70.
  4. Caucheteux, Charlotte, and Jean-Rémi King. 2022. “Brains and Algorithms Partially Converge in Natural Language Processing.” Communications Biology 5 (134).
  5. Clark, Andy. 2013. “Whatever Next? Predictive Brains, Situated Agents, and the Future of Cognitive Science.” Behavioral and Brain Sciences 36 (3): 181–204.
  6. Elkins, Katherine. 2022. The Shapes of Stories: Sentiment Analysis for Narratives. Cambridge University Press.
  7. Evert, Stefan, Thomas Proisl, Frederike Jannidis, et al. 2017. Understanding and Explaining Delta Measures for Authorship Attribution. Digital Scholarship in the Humanities 32:4–16.
  8. Herrmann, J. Berenike, Karina van Dalen-Oskam, and Christof Schöch. 2015. Revisiting Style, a Key Concept in Literary Studies. Journal of Literary Theory 9 (1): 25–52.
  9. Hogan, Patrick Colm. 2016. Beauty and Sublimity: A Cognitive Aesthetics of Literature and the Arts. Cambridge University Press.
  10. Holtzman, Ari, Jan Buys, Li Du, et al. 2020. The Curious Case of Neural Text Degeneration. In Proceedings of the International Conference on Learning Representations (ICLR).
  11. Jelinek, F., R. L. Mercer, L. R. Bahl, et al. 1977. Perplexity—A Measure of the Difficulty of Speech Recognition Tasks. The Journal of the Acoustical Society of America 62 (S1).
  12. Ji, Fengyuan. 2004. Linguistic Engineering: Language and Politics in Mao’s China. Honolulu: University of Hawai’i Press.
  13. Kirchenbauer, John, Jonas Geiping, Yuxin Wen, et al. 2023. A Watermark for Large Language Models. In Proceedings of the 40th International Conference on Machine Learning, edited by Andreas Krause, Emma Brunskill, Kyunghyun Cho, et al., 202:17061–17084. Proceedings of Machine Learning Research. PMLR, 23–29 Jul.
  14. Kumar, Sreejan, Theodore R. Sumers, Takateru Yamakoshi, et al. 2024. “Shared Functional Specialization in Transformer-Based Language Models and the Human Brain”. Nature Communications 15 (5523).
  15. Kurzynski, Maciej. 2023. The Stylometry of Maoism: Quantifying the Language of Mao Zedong. In Proceedings of the Joint 3rd International Conference on Natural Language Processing for Digital Humanities and 8th International Workshop on Computational Linguistics for Uralic Languages, 76–81. Association for Computational Linguistics. https://aclanthology.org/2023.nlp4dh-1.9/.
  16. Kurzynski, Maciej. 2024. Perplexity Games: Maoism vs. Literature through the Lens of Cognitive Stylometry. Journal of Data Mining and Digital Humanities NLP4DH:23. https://doi.org/10.46298/jdmdh.13131.
  17. Leese, Daniel. 2011. Mao Cult: Rhetoric and Ritual in China’s Cultural Revolution. Cambridge University Press.
  18. Li, Tuo. 1998. 汪曾祺与现代汉语写作——兼谈毛文体 [Wang Zengqi and Modern Chinese Writing: Also on the Style of Mao’s Writings]. 花城 [Huacheng].
  19. Link, Perry. 2013. An Anatomy of Chinese: Rhythm, Metaphor, Politics. Harvard University Press.
  20. Lutosławski, Wincenty. 1898. Principes de Stylométrie Appliqués à la Chronologie des Oeuvres de Platon. Revue des Études Grecques 11 (41): 61–81.
  21. Markov, Andrey. 2006. “An Example of Statistical Investigation of the Text Eugene Onegin Concerning the Connection of Samples in Chains.” Science in Context 19 (4): 591–600.
  22. McDonald, Scott A., and Richard C. Shillock. 2004. “Lexical Predictability Effects on Eye Fixations During Reading.” In The On-line Study of Sentence Comprehension: Eyetracking, ERPs and Beyond, edited by Manuel Carreiras and Charles Clifton Jr., 77–94. Psychology Press.
  23. Mitchell, Eric, Yoonho Lee, Alexander Khazatsky, et al. 2023. “Detect-GPT: Zero-Shot Machine-Generated Text Detection Using Probability Curvature.” In Proceedings of the 40th International Conference on Machine Learning. ICML’23. Honolulu, Hawaii, USA.
  24. Mosteller, Frederick, and David L. Wallace. 1963. “Inference in an Authorship Problem.” Journal of the American Statistical Association 58 (302): 275–309.
  25. Reagan, Andrew J, Lewis Mitchell, Dilan Kiley, et al. 2016. “The Emotional Arcs of Stories are Dominated by Six Basic Shapes.” EPJ Data Science 5:31.
  26. Rybicki, Jan, Maciej Eder, and David L. Hoover. 2016. “Computational Stylistics and Text Analysis.” In Doing Digital Humanities, edited by Constance Crompton, Richard J. Lane, and Ray Siemens, 123–144. Routledge.
  27. Schoenhals, Michael. 2007. “Demonising Discourse in Mao Zedong’s China: People vs Non‐People.” Totalitarian Movements and Political Religions 8 (3): 465– 482.
  28. Shannon, Claude E. 1951. “Prediction and Entropy of Printed English.” Bell System Technical Journal 30 (1): 50–64.
  29. Shklovsky, Viktor. 2017. Viktor Shklovsky: A Reader. Edited and translated by Alexandra Berlina. Bloomsbury Academic.
  30. Snow, Don. 2004. Cantonese as Written Language: The Growth of a Written Chinese Vernacular. Hong Kong University Press.
  31. Tikochinski, Rafael, Ariel Goldstein, Yoav Meiri, et al. 2025. “Incremental accumulation of linguistic context in artificial and biological neural networks,” Nature Communications 16:803.
  32. Tsur, Reuven. 2008. Toward a Theory of Cognitive Poetics. Brighton: Sussex Academic Press.
  33. Varnum, Michael E. W., Nicolas Baumard, Mohammad Atari, and Kurt Gray. 2024. “Large Language Models Based on Historical Text Could Offer Informative Tools for Behavioral Science.” Proceedings of the National Academy of Sciences 121 (42): e2407639121. https://doi.org/10.1073/pnas.2407639121.
  34. Wu, Yaru, Yuri Bizzoni, Pascale Moreira, et al. 2024. “Perplexing Canon: A Study on GPT-Based Perplexity of Canonical and Non-Canonical Literary Works,” Proceedings of the 8th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2024), 172–184. 
  35. Yu, Yijiong, Ziyun Dai, Zekun Wang, et al. 2025. OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training. arXiv preprint arXiv:2501.08197.
  36. Zhao, Xuandong, Prabhanjan Ananth, Lei Li, et al. 2023. Provable Robust Watermarking for AI-Generated Text. arXiv preprint arXiv:2306.17439.