DH 2026

Daejeon, July 27–31

Wed, July 2916:30–18:00S054209-211
Short Paper

Beyond Standard Language: Detecting Rule-Changing Creativity in Experimental Literature.

Silvia Lilli
Independent Researcher · silvialilli@hotmail.it

NLP tools have enormously increased our possibilities in exploring verbal human expression, allowing us to automatize demanding tasks such as lemmatization, morphological and syntactic tagging, Named Entity Recognition (NER), word sense disambiguation, and so on. However, they are built essentially on and for high-resources languages. This produces a significant gap, as better-represented languages end up having at their disposal a huge set of instruments that continue to increase their knowledge at the expense of less represented languages (Hovy / Spruit 2016).

This is a well-known issue (Pakray et al. 2025), which has focused on the confrontation between different languages. Nevertheless, similar problems are to be encountered within a language itself, when we exit the elusive boundaries of standard varieties and face alterations the machine cannot cope with. While on diatopic, chronological and sociolinguistic varieties NLP resources still possess margins of application given their social sharing and multiple attestations, the question changes when we deal with idiolectic varieties, such as those created by authors for literary purposes. Both low-resource languages and experimental literary idiolects challenge the fundamental assumption of ‘standardness’ underlying most NLP tools (Plank 2016), requiring adapted methodological approaches (Zampieri et al. 2020).

As Chomsky (1964, 1966) and other scholars (Sampson 2016; Bergs 2019) have pointed out, linguistic mutations happen along two directions: within the codified system – the creativity everybody performs while speaking or writing, particularly marked for artistic purposes – and beyond it – the creativity which allows for language evolution at all. The latter kind, variously defined as ‘rule-changing’ (Chomsky 1964: 22), ‘rule-breaking’ (Lecercle 2017), or lately ‘E-creativity’ (meaning ‘enlarging creativity’, Sampson 2016: 17) has been used by literary authors for expressive purpose, and also to protest the normative (political and social) institution represented by language itself. Experimental writers, despite being a minority, can be found sparsely across literatures and times: e.g., Rabelais in the Renaissance, Lewis Carroll in the Victorian Age, James Joyce or Carlo Emilio Gadda in the Modernist era. Their style results in a mixture of diverse components, derived from different languages (i.e. plurilingualism), or different context of use (archaisms, colloquialisms), sometimes with coinage of entirely new words or alteration of existing forms and allowable structures. Despite the seemingly chaotic construction, each idiolectal language maintains an internal coherence which rules the deformation, necessary to ensure communicability of authorial expression towards readers (Den Ouden 1975: 18, Sanguineti 1999).

Therefore, attempts at applying automatic text analysis can rely on two complementary assumptions:

  • Experimental texts willingly deviate from standard word-forms or syntactic structures.
  • Experimental texts show consistent deformation patterns within a single work or group of works of the same author.

Both these criteria allow for development of automated processing, to highlight the linguistic features which characterize an author’s experimentation and connect the stylistic layer to the semantic and thematic one, in a fully integrated perspective between language and meaning, and quantitative and qualitative analysis (Herrmann et al. 2015, Herrmann 2017, Herrmann et al. 2021). “Engaging” with the literary text requires developing methodologies which maintain a high degree of interpretability, as deviations are not dischargeable noise but significance triggers (Lilli 2025).

This presentation proposes two complementary computational approaches to experimental literary texts, integrating traditional text query methods with language model resources:

  • Bottom-up approach: pattern-based detection. Building regular expressions matching deviation patterns previously identified through close reading – i.e. phonetic, morphological, lexical, or syntactic elements which contribute to distancing the text from standard language – then filtering matches by confronting them with existing dictionaries or lemmatizers. This method takes advantage of resources for standard varieties to identify, by negation, unmatched forms or structures which may signal markedness. While effective for known deviation types, this method requires prior identification of experimental strategies through close reading, thus limiting its scalability and capacity for discovering unanticipated patterns. A similar unsupervised approach has also been adopted by Sluyter-Gäthje and Trilcke (2022) and Nini (2025).

I have tested this approach on two 20th century Italian experimental authors (Giovanni Testori and Stefano D’Arrigo), achieving good results in detecting authorial deviations. For D’Arrigo’s corpus, I focused on two specific features: intensifying reduplications (e.g. “mare mare”) and parasynthetic verbal neologisms modelled on Sicilian dialect (e.g. “appiccionarsi”), demonstrating how automated retrieval procedures proved more efficient and reliable than manual search. For Testori’s corpus, combining automatic and manual annotation, I systematically explored stylistically marked and deviant elements according to their source of derivation (dialect, archaic or literary register, foul language, authorial innovations) and single features. This extensive quantification of the author’s style yielded fundamental insights into his experimentation (e.g., the correlation observed between different features illuminated the pragmatic and semantic functions of deviations; see Lilli 2025), and allowed me to overcome impressionistic assumptions (such as the overestimated presence of obscene language, far outnumbered by Latin adaptations).

  • Top-down approach: information-theoretic measures. Applying measures of ‘surprisal’ and ‘entropy’ from information theory (Shannon 1948; applied to linguistics by Hale 2016): calculating through Language Models the expectation based on probability and the quantity of novelty in given information. The combination of these measures (Degaetano-Ortlieb / Teich 2017, Slaats / Martin 2025) can be interpreted as detector for marked forms, especially when high surprisal values coincide with low entropy contexts (thus, when unexpected words occur in highly predictable sequences). To enrich the interpretability of these peaks, surprisal can be paired with additional linguistic features such as part-of-speech (POS) distribution, syntactic dependencies, and semantic categories. This approach operates independently of pre-existing taxonomies of experimental features, potentially revealing deviation patterns not anticipated through traditional close reading.

This method is currently under development and promises broader applicability beyond experimental literary languages to poetic language in general. While it has been often applied in psychology (de Varda / Marelli 2022, Giulianelli et al. 2023, Meister et al. 2024, Staub 2025), its potential in Literary Studies is only beginning to be explored (Kontoyiannis 1997, Kozhemyakina et al. 2023, Zhang / Liang 2024).

These two methods create a virtuous cycle: the top-down approach identifies zones of high markedness not predicted by existing taxonomies, while the bottom-up method refines understanding of how specific deviations function within the author’s idiolectic system. This mixed methodology demonstrates how computational analysis and literary interpretation can integrate iteratively, maintaining the interpretability necessary for meaningful engagement with experimental texts while expanding the scale and scope of analysis.

References
  1. Bergs, Alexander (2019): “What, If Anything, Is Linguistic Creativity?”, in: Gestalt Theory 41, 2: 17383. DOI: 10.2478/gth-2019-0017.
  2. Chomsky, Noam (1966): Cartesian Linguistics. New York and London: Harper & Row.
  3. Chomsky, Noam (1964): Current Issues in Linguistic Theory. The Hauge: Mouton & Co.
  4. de Varda, Andrea / Marelli, Marco (2022): “The Effects of Surprisal across Languages: Results from Native and Non-Native Reading”, in: He, Yulan / Ji, Heng / Li, Sujian / Liu, Yang / Chang, Chua-Hui (eds.): Findings of the Association for Computational Linguistics, AACL-IJCNLP 2022, Association for Computational Linguistics 138–144. DOI: 10.18653/v1/2022.findings-aacl.13.
  5. Degaetano-Ortlieb, Stefania / Teich, Elke (2017): “Modeling Intra-Textual Variation with Entropy and Surprisal: Topical vs. Stylistic Patterns”, in: Alex, Beatrice / Degaetano-Ortlieb, Stefania / Feldman, Anna / Kazantseva, Anna / Reiter, Nils / Szpakowicz, Stan (eds.): Proceedings of the Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, Vancouver, Canada: Association for Computational Linguistics 68–77. DOI: 10.18653/v1/W17-2209.
  6. Den Ouden, Bernard D. (1975): Language and Creativity. An Interdisciplinary Essay in Chomskyan Humanism. Berlin, Boston: De Gruyter Mouton. DOI: 10.1515/9783110883473.
  7. Giulianelli, Mario / Wallbridge, Sarenne / Fernández, Raquel (2023): “Information Value: Measuring Utterance Predictability as Distance from Plausible Alternatives”, in: Bouamor, Houda / Pino, Juan / Bali, Kalika (eds): Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, December 2023, Association for Computational Linguistics 5633–5653. DOI: 10.18653/v1/2023.emnlp-main.343.
  8. Hale, John (2016): “Information-Theoretical Complexity Metrics”, in: Language and Linguistics Compass 10, 9: 397–412. DOI: 10.1111/lnc3.12196.
  9. Herrmann, Berenike J. / van Dalen-Oskam, Karina / Schöch, Christof (2015): “Revisiting Style, a Key Concept in Literary Studies”, in: Journal of Literary Theory 9, 1: 25–52. DOI: 10.1515/jlt-2015-0003.
  10. Herrmann, J. Berenike / Jacobs, Arthur M. / Piper, Andrew (2021): “Computational Stylistics”, in: Kuiken, Donald / Jacobs, Arthur M. (eds): Handbook of Empirical Literary Studies, Berlin, Boston: De Gruyter 451–486. DOI: 10.1515/9783110645958-018.
  11. Herrmann, J. Berenike (2017): “In a Test Bed with Kafka. Introducing a Mixed-Method Approach to Digital Stylistics”, in: Digital Humanities Quarterly 11, 4 <http://www.digitalhumanities.org/dhq/vol/11/4/000341/000341.html> [29.04.2026].
  12. Hovy, Dirk / Spruit, Shannon L (2016): “The Social Impact of Natural Language Processing”, in:  Erk, Katrin / Smith, Noah A., Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Volume 2: Short Papers, Berlin, Germany: Association for Computational Linguistics 591–598. DOI: 10.18653/v1/P16-2096.
  13. Kontoyiannis, Ioannis (1997): “The Complexity and Entropy of Literary Styles”, NSF Technical Report 97. Department of Statistics, Stanford University, <http://pages.cs.aueb.gr/~yiannisk/PAPERS/english.pdf>.
  14. Kozhemyakina, Olga / Barakhnin, Vladimir / Shashok, Natalia / Kozhemyakina, Elina (2023). ‘The Question of Studying Information Entropy in Poetic Texts“, in: Applied Sciences 13, 20: 11247. DOI: 10.3390/app132011247.
  15. Lecercle Jean-Jacques (2017): “Linguistic Creativity: Rule-Governed or Rule-Breaking?”, in: Recherches anglaises et nord-américaines, n. 50: Discourse, Boundaries and Genres in English Studies: an Assessment: 69–80. DOI: 10.3406/ranam.2017.1549.
  16. Lilli, Silvia (2025): “Creativity, Invention and Linguistic Analysis: A Case Study”, in: Umanistica Digitale 9, 19: 105–124. DOI: 10.6092/issn.2532-8816/21005.
  17. Meister, Clara / Giulianelli, Mario / Pimentel, Tiago (2024): “Towards a Similarity-Adjusted Surprisal Theory”. arXiv. DOI: 10.48550/arXiv.2410.17676.
  18. Nini, Andrea (2025): “Examining an author’s individual grammar”, in: Comparative Literature Goes Digital Workshop, Digital Humanities 2025, Universidade Nova de Lisboa, Lisbon, Portugal, July 2025. DOI: 10.5281/zenodo.15971103.
  19. Pakray, Partha / Gelbukh, Alexander / Bandyopadhyay, Sivaji (2025): “Natural Language Processing Applications for Low-Resource Languages”, in: Natural Language Processing 31, 2: 183–97. DOI: 10.1017/nlp.2024.33.
  20. Plank, Barbara (2016): “What to Do about Non-Standard (or Non-Canonical) Language in NLP”, in: Proceedings of the 13th Conference on Natural Language Processing, KONVENS 2016, arXiv. DOI: 10.48550/arXiv.1608.07836.
  21. Sampson, Geoffrey (2016): “Two Ideas of Creativity”, in: Hinton, Martin (ed.): Evidence, Experiment and Argument in Linguistics and the Philosophy of Language, Bern: Peter Lang 15–20.
  22. Sanguineti, Edoardo (1999): “Il plurilinguismo nelle scritture novecentesche”, in: Sertoli, Giuseppe / Miglietta, Goffredo (eds.): Transiti letterari e culturali, vol. 1, Trieste: Edizioni Università di Trieste 17–31.
  23. Shannon, Claude E. (1948): “A Mathematical Theory of Communication”, in: Bell System Technical Journal 27, 3: 379–423. DOI: 10.1002/j.1538-7305.1948.tb01338.x.
  24. Slaats, Sophie, / Martin, Andrea E (2025): “What’s Surprising About Surprisal”, in: Computational Brain & Behavior 8: 233–248. DOI: 10.1007/s42113-025-00237-9.
  25. Sluyter-Gäthje, Henny / Trilcke, Peer (2022): “Poetry as Error. A ‘Tool Misuse’ Experiment on the Processing of German Language Poetry”, ADHO 2022, Tokyo <https://dh-abstracts.library.cmu.edu/works/11743> [29.04.2026].
  26. Staub, Adrian (2025): “Predictability in Language Comprehension: Prospects and Problems for Surprisal”, in: Annual Review of Linguistics 11: 17–34. DOI: 10.1146/annurev-linguistics-011724-121517.
  27. Zampieri, Marcos / Nakov, Preslav / Scherrer, Yves (2020): “Natural Language Processing for Similar Languages, Varieties, and Dialects: A Survey”, in Natural Language Engineering 26, 6: 595–612. DOI: 10.1017/S1351324920000492.
  28. Zhang, Tianyi / Liang, Junying (2024): “Quantifying the Information Flow of Long Narratives: A Case Study of Jane Austin’s Works”, in: Journal of Theory and Practice in Humanities and Social Sciences 1, 3: 7–12 <https://woodyinternational.com/index.php/jtphss/article/view/28>.