Daejeon, July 27–31
NLP tools have enormously increased our possibilities in exploring verbal human expression, allowing us to automatize demanding tasks such as lemmatization, morphological and syntactic tagging, Named Entity Recognition (NER), word sense disambiguation, and so on. However, they are built essentially on and for high-resources languages. This produces a significant gap, as better-represented languages end up having at their disposal a huge set of instruments that continue to increase their knowledge at the expense of less represented languages (Hovy / Spruit 2016).
This is a well-known issue (Pakray et al. 2025), which has focused on the confrontation between different languages. Nevertheless, similar problems are to be encountered within a language itself, when we exit the elusive boundaries of standard varieties and face alterations the machine cannot cope with. While on diatopic, chronological and sociolinguistic varieties NLP resources still possess margins of application given their social sharing and multiple attestations, the question changes when we deal with idiolectic varieties, such as those created by authors for literary purposes. Both low-resource languages and experimental literary idiolects challenge the fundamental assumption of ‘standardness’ underlying most NLP tools (Plank 2016), requiring adapted methodological approaches (Zampieri et al. 2020).
As Chomsky (1964, 1966) and other scholars (Sampson 2016; Bergs 2019) have pointed out, linguistic mutations happen along two directions: within the codified system – the creativity everybody performs while speaking or writing, particularly marked for artistic purposes – and beyond it – the creativity which allows for language evolution at all. The latter kind, variously defined as ‘rule-changing’ (Chomsky 1964: 22), ‘rule-breaking’ (Lecercle 2017), or lately ‘E-creativity’ (meaning ‘enlarging creativity’, Sampson 2016: 17) has been used by literary authors for expressive purpose, and also to protest the normative (political and social) institution represented by language itself. Experimental writers, despite being a minority, can be found sparsely across literatures and times: e.g., Rabelais in the Renaissance, Lewis Carroll in the Victorian Age, James Joyce or Carlo Emilio Gadda in the Modernist era. Their style results in a mixture of diverse components, derived from different languages (i.e. plurilingualism), or different context of use (archaisms, colloquialisms), sometimes with coinage of entirely new words or alteration of existing forms and allowable structures. Despite the seemingly chaotic construction, each idiolectal language maintains an internal coherence which rules the deformation, necessary to ensure communicability of authorial expression towards readers (Den Ouden 1975: 18, Sanguineti 1999).
Therefore, attempts at applying automatic text analysis can rely on two complementary assumptions:
Both these criteria allow for development of automated processing, to highlight the linguistic features which characterize an author’s experimentation and connect the stylistic layer to the semantic and thematic one, in a fully integrated perspective between language and meaning, and quantitative and qualitative analysis (Herrmann et al. 2015, Herrmann 2017, Herrmann et al. 2021). “Engaging” with the literary text requires developing methodologies which maintain a high degree of interpretability, as deviations are not dischargeable noise but significance triggers (Lilli 2025).
This presentation proposes two complementary computational approaches to experimental literary texts, integrating traditional text query methods with language model resources:
I have tested this approach on two 20th century Italian experimental authors (Giovanni Testori and Stefano D’Arrigo), achieving good results in detecting authorial deviations. For D’Arrigo’s corpus, I focused on two specific features: intensifying reduplications (e.g. “mare mare”) and parasynthetic verbal neologisms modelled on Sicilian dialect (e.g. “appiccionarsi”), demonstrating how automated retrieval procedures proved more efficient and reliable than manual search. For Testori’s corpus, combining automatic and manual annotation, I systematically explored stylistically marked and deviant elements according to their source of derivation (dialect, archaic or literary register, foul language, authorial innovations) and single features. This extensive quantification of the author’s style yielded fundamental insights into his experimentation (e.g., the correlation observed between different features illuminated the pragmatic and semantic functions of deviations; see Lilli 2025), and allowed me to overcome impressionistic assumptions (such as the overestimated presence of obscene language, far outnumbered by Latin adaptations).
This method is currently under development and promises broader applicability beyond experimental literary languages to poetic language in general. While it has been often applied in psychology (de Varda / Marelli 2022, Giulianelli et al. 2023, Meister et al. 2024, Staub 2025), its potential in Literary Studies is only beginning to be explored (Kontoyiannis 1997, Kozhemyakina et al. 2023, Zhang / Liang 2024).
These two methods create a virtuous cycle: the top-down approach identifies zones of high markedness not predicted by existing taxonomies, while the bottom-up method refines understanding of how specific deviations function within the author’s idiolectic system. This mixed methodology demonstrates how computational analysis and literary interpretation can integrate iteratively, maintaining the interpretability necessary for meaningful engagement with experimental texts while expanding the scale and scope of analysis.