DH 2026

Daejeon, July 27–31

Short Paper

Modeling Intra-Author Lexical Variation in Baldomero Lillo: A Computational Approach Based on Corpus StratificationA Computational Approach Based on Corpus Stratification

María Teresa Filipigh
National University of Distance Education, Spain · mfilipigh1@alumno.uned.es

INTRODUCTION

In the nineteenth century and the early decades of the twentieth, newspapers and books coexisted as the principal vehicles for literary circulation in Latin America, fostering new forms of sociability and authorial professionalization (Botinelli 2023, Rama 1998, Rivera, 1998). Within this ecosystem, the short story often first appeared in periodicals and was later incorporated –frequently after substantial revision – into single-author volumes, generating divergent textual states that complicate philological reconstruction and stylistic analysis.

Baldomero Lillo (1867–1923, Chile) offers a representative case: dozens of his stories circulated in Chilean newspapers and magazines, a curated subset –selected and revised by Lillo– entered Sub terra (1904, 1917) and Sub sole (1907), while additional texts surfaced posthumously without confirmed revision (Blanco 2022). While Lillo has been examined through social naturalism and mining representation (e.g. Álvarez 2010; Concha 2008, Río, 2013), the stylistic and semantic variation generated by this fragmented transmission has received little attention. Although intra-author stylometric analyses of Hispanic writers have gained visibility in recent years (e.g. Boto Bravo 2018, Celma / Ruiz 2021, Hernández-Lorenzo 2023), applications to the Spanish American short story remain almost non-existent (Kovalev 2024).

The present short paper proposal adopts a computational lexical–semantic approach, grounded in philological modularization, to examine intra-author variation in Lillo’s short stories. Its objectives are: (i) to model the lexical and semantic properties of three corpus strata –press, book-revised, and posthumous stories–; and (ii) to determine how publication medium, revision practices, and transmission history shape Lillo’s internal stylistic variation. This proposal forms part of an ongoing doctoral dissertation on Lillo's style, which provides the groundwork necessary to integrate this stratified corpus design into a broader analysis of his narrative practice within computational stylistics.

CORPUS TRANSMISSION AND CORPUS STRATIFICATION

Lillo’s fiction survives in divergent textual states. Press versions frequently differ from book versions; anthologized stories constitute the only authorially stabilized texts; other stories survive solely in posthumous versions, and recent archival discoveries continue to expand the corpus (Álvarez / Bello Maldonado 2008, Blanco Casals 2022). Treating all texts as a homogeneous dataset would collapse these differences into an undifferentiated stylistic average.

To avoid this, the corpus is divided into three philologically grounded strata: (a) Corpus_prensa (24 files; 77,954 tokens), consisting of texts only attested in newspapers and magazines but not included in the anthologies; (b) Corpus_libro (24 files; 73,376 tokens), comprising texts authorially revised or prepared for inclusion in the anthologies; (c) Corpus_postuma (2 files; 3,532 tokens), representing texts without demonstrable authorial revision. This stratified corpus design provides a systematic framework for analyzing lexical cohesion, semantic specialization, and distributional patterns across the different stages of Lillo’s textual transmission.

METHODOLOGY

All stories were retrieved from the critical edition by Álvarez and Bello Maldonado (2008), converted into plain text, lemmatized with FreeLing (Pradó / Stanilovsky 2012), and processed in RStudio; network visualizations were produced in Gephi. The analysis combines a stylometric baseline with lexical-semantic techniques adapted to the stratified structure of the corpus. In line with intra-author stylometric work that applies PCA and Delta to model internal variation within a single oeuvre (Calvo Tello 2019; Kovalev 2026), a global PCA was performed to establish Lillo’s overall stylistic baseline. However, MFW–based stylometry alone is not sensitive to contextual and medium-dependent variation and therefore requires complementary approaches. Accordingly, the analysis focuses on lexical distribution, tf-idf distinctiveness, co-occurrence structures, and length-corrected measures of lexical density (Johansson 2008). Co-occurrence graphs are built from the most frequent lexical lemmas in each corpus, ranked by relative frequency, using sentence boundaries as context windows (Rojo 2021) and logDice as the association measure (Rychlý 2008); only edges above a minimum raw co-occurrence threshold are retained for visualization. Sentiment analysis provides an approximate tonal mapping, interpreted cautiously due to limitations in historical Spanish. These procedures align with established methods in computational literary studies and corpus linguistics (Biber 1995; Capsada / Torruella 2017; Eder et al. 2016; Jockers 2013; Schöch 2018).

RESULTS: STYLISTIC VARIATION AND THEMATIC COHESION ACROSS CORPUS STRATA

The global PCA (Fig. 1) provides an initial overview of the corpus. In the complete set of fifty stories (Fig. 1A), most texts form a compact cloud, suggesting a shared stylistic profile; however, the shortest narratives –“La carga”, “El oro” and “[Cuento sin título]”– occupy peripheral positions. For this reason, the PCA of the full corpus is treated as exploratory. When texts below 1,500 tokens are excluded (Fig. 1B), the remaining forty-one stories retain a broad central concentration. This threshold is supported by a marked gap in the length distribution between “Cambiadores” (1,416 tokens) and “Las nieves” (1,881). The most distant cases –“Víspera de difuntos”, “Carlitos” and “El anillo”– are individual texts rather than a clear chronological sequence, indicating that chronology alone does not account for the corpus’s internal variation.

Clear contrasts become apparent when examining specific lexical and semantic layers. Lexical density values corroborate these differences: although the overall figures are nearly identical Corpus_prensa (0.509), whereas Corpus_libro (0.511), the latter displays greater internal stability, whereas the press corpus shows slightly higher dispersion, consistent with its more heterogeneous thematic and register profile. Corpus_postuma, represented by only two stories, is interpreted with caution because of its limited size.

ABFigure 1. PCA of Baldomero Lillo’s short stories. (A) Exploratory PCA of all fifty stories. (B) PCA restricted to forty-one stories exceeding 1,500 tokens. Labels indicate year of first publication and title of each story

Figure 2. Lemma–lemma co-occurrence graph of Corpus_prensa, showing the heterogeneous semantic organization of the press corpusCorpus_prensa displays the highest internal heterogeneity. Its lexical distribution featuring Chilean regionalisms, colloquialisms, and vocabulary associated with urban and maritime settings. Co-occurrence networks (Fig. 2) reveal diverse semantic domains – occupations, urban life, costumbrista scenes– and a more balanced emotional polarity. This stratum reflects writing adapted to a newspaper audience and more open to thematic experimentation.

Figure 3. Lemma–lemma co-occurrence graph of Corpus_libro, illustrating the more concentrated semantic organization of the revised anthology corpus

By contrast, Corpus_libro exhibits greater lexical cohesion and a more concentrated semantic profile. Mining terminology, rural vocabulary, and lexicon related to bodily perception and labor form dense, recurrent associations. Co-occurrence graphs (Fig. 3) reveal a strongly interconnected core, while regionalisms are reduced relative to the press stories, suggesting an authorial tendency toward lexical universalization in the stabilized anthologies. Negative emotional polarity predominates, consistent with the thematic orientation of Sub terra and Sub sole.

Finally, Corpus_postuma, though limited in size, introduces unexpected registers, such as humor in “Inamible” and light sentimentalism in “[Cuento sin título]”. Figure 4 reinforces this contrast: the latter combines unusually high lexical density with extreme brevity, standing apart from the short narratives in the corpus, which tend to show lower density values. Together with its lack of title, this profile may point a condensed or unfinished draft. “Inamible”, by contrast, falls within the general density-length distribution of the corpus, a profile consistent with a more advanced stage of textual elaboration. In the exploratory PCA (see Fig. 1), both lexical profiles are integrated into overall stylistic configuration: the shorter text follows the peripheral pattern also observed in other very short narratives, whereas the other one remains part of the broader configuration of the length-controlled analysis. This suggests that deeper stylistic features remain stable even in texts not clearly intended for publication. Variation arises primarily at the thematic rather than structural level.

Figure 4. Lexical density and text length in the complete corpus. The two stories in Corpus Postuma are labelled for comparison

CONCLUSIONS

The computational analyses employed in this study disclose meaningful intra-author variation in Baldomero Lillo’s fiction. Although global stylometry confirms a cohesive authorial signal across the corpus, its stratification into press, book, and posthumous materials reveal differentiated lexical, thematic, and dialectal configurations associated with distinct modes of circulation and textual transmission.

The press-only stories display a more diverse lexical and thematic repertoire, together with a higher concentration of regionalisms and dialectal features. That profile is more closely connected to issues and cultural references circulating in the contemporary press, suggesting an authorial positioning attuned to a newspaper readership. The book corpus, by contrast, nuances his exclusive association with mining fiction, reinforces the social dimension of his work. In turn, its lower degree of dialectal marking, combined with its internal stylistic configuration, is consistent with a process of stabilization associated with revision and republication. Despite its limited size, the posthumous corpus further extends this profile by preserving registers that are less prominent in the other strata.

More broadly, this study demonstrates that combining philological grounded corpus stratification with computational analysis provides a replicable framework for Latin American authors whose works circulated across heterogeneous media. As part of an ongoing doctoral dissertation, this approach establishes a foundation for expanding the analysis to syntactic, morphological, and dialectal dimensions, demonstrating that the present proposal constitutes one component of a larger and more comprehensive research program.

References
  1. Álvarez Arenas, Ignacio / Bello Maldonado, Hugo (eds.) (2008): Baldomero Lillo. Obra completa. Santiago de Chile: Ediciones Universidad Alberto Hurtado.
  2. Álvarez, Ignacio (2010): “Huérfanos y mineros: notas para una evaluación de la estrategia representativa del obrero en los cuentos de Baldomero Lillo”, in: Anales de Literatura Chilena 14: 93–116.
  3. Blanco Casals, Marcelo Javier (2022): “El hallazgo de un cuento olvidado de Baldomero Lillo”, in: Anales de Literatura Chilena 37: 243–247 https://doi.org/10.7764/ANALESLITCHI.37.18 [28.04.2026].
  4. Boto Bravo, María Ángeles (2018): “Mapa estilométrico de la narrativa de Eduardo Mendoza: aproximación a un análisis estilístico computacional de textos literarios”, in: Epos: Revista de Filología 33: 99–114 https://doi.org/10.5944/epos.33.2017.18281 [28.04.2026].
  5. Bottinelli Wolleter, Mónica / Sanhueza, María (2023): “Literatura, prensa y mercado en el siglo XIX latinoamericano: dislocaciones de la hegemonía letrada”, in: Atenea 528: 7–14 https://www.scielo.cl/scielo.php?pid=S0716-58112023000100013&script=sci_arttext [28.04.2026].
  6. Calvo Tello, José (2019): “Delta inside Valle-Inclán: stylometric classification of periods and groups of his novels”, in: Romanische Studien 6: 151–163.
  7. Capsada Blanch, Ramon / Torruella Casañas, Joan (2017): “Métodos para medir la riqueza léxica de los textos: revisión y propuesta”, in: Verba 44: 347–408 https://doi.org/10.15304/verba.44.3155 [28.04.2026].
  8. Celma Valero, María Pilar / Ruiz Urbón, Cristina (2021): “Una aproximación a la escritura de Miguel Delibes desde la estilometría”, in: Tonos Digital 41: 1–13.
  9. Concha, Jaime (2008): “Lillo y los condenados de la tierra”, in: Álvarez Arenas, Ignacio / Bello Maldonado, Hugo (eds.): Baldomero Lillo. Obra completa. Santiago de Chile: Ediciones Universidad Alberto Hurtado 14–58.
  10. Eder, Maciej / Rybicki, Jan / Kestemont, Mike (2016): “Stylometry with R: a package for computational text analysis”, in: The R Journal 8, 1: 107–121.
  11. Guerrero del Río, Eduardo (2013): “Baldomero Lillo: padre del realismo social”, in: Mensaje 62, 621: 56–58 https://dialnet.unirioja.es/servlet/articulo?codigo=4378997 [28.04.2026].
  12. Hernández-Lorenzo, Laura (2023): “Una primera aproximación a los géneros literarios de la obra cervantina desde la estilometría y el análisis de redes”, in: Tirant 26: 175–192 https://doi.org/10.7203/tirant.26.27873 [28.04.2026].
  13. Hoover, David L. (2007): “Quantitative analysis and literary studies”, in: Siemens, Ray / Schreibman, Susan (eds.): A Companion to Digital Literary Studies. Oxford: Blackwell 517–533.
  14. Jockers, Matthew L. (2013): Macroanalysis: Digital Methods and Literary History. Urbana / Chicago / Springfield: University of Illinois Press.
  15. Johansson, Victoria (2008): “Lexical diversity and lexical density in speech and writing: a developmental perspective”, in: Working Papers in Linguistics 53: 61–79.
  16. Kovalev, Boris (2024): “The birth of the third author: stylometric analysis of the stories of Honorio Bustos Domecq”, in: Digital Scholarship in the Humanities 39, 1: 112–130 https://doi.org/10.1093/llc/fqad067 [28.04.2026].
  17. Kovalev, Boris (2026): “La Periodización de las Novelas de Gabriel García Márquez con los Métodos Estilométricos”, in: Linguistics & Polyglot Studies 12, 1: 103–117 https://doi.org/10.24833/2410-2423-2026-1-46-103-117 [28.04.2026].
  18. Padró, Lluís / Stanilovsky, Evgeny (2012): “FreeLing 3.0: towards wider multilinguality”, in: Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC 2012). Istanbul: European Language Resources Association (ELRA) 2473–2479.
  19. Rama, Ángel (1998): La ciudad letrada. Montevideo: Arca.
  20. Rychlý, Pavel (2008): “A lexicographer-friendly association score”, in: Sojka, Petr / Horák, Aleš (eds.): Proceedings of RASLAN 2008: Recent Advances in Slavonic Natural Language Processing. Brno: Masaryk University 6–9 https://www.sketchengine.eu/wp-content/uploads/2015/03/Lexicographer-Friendly_2008.pdf [28.04.2026].
  21. Rivera, Jorge (1998): El escritor y la industria cultural. Fondo de Cultura Económica.
  22. Rojo, Guillermo (2021): Introducción a la lingüística de corpus en español. London / New York: Routledge.