DH 2026

Daejeon, July 27–31

Thu, July 3016:30–18:00S050104
Short Paper

Mapping the Nation: A Lexical Network Analysis of Argentine Identity (1830–1910)

María Teresa Filipigh
National University of Distance Education, Spain · mfilipigh1@alumno.uned.es
María Gimena Del Río Riande
Consejo Nacional de Investigaciones Científicas y Técnicas · gdelrio@conicet.gov.ar

INTRODUCTION

This short paper proposal presents an ongoing study that applies macroanalysis approaches (Jockers 2013) and distant reading methods (Jänicke 2016, Jänicke et al.  2015) to the discursive construction of Argentine national identity between 1830 and 1910 (Martínez Estrada 1933a, 1940b). We work with a heterogeneous corpus –including diaries, memoirs, travelogues, political essays, narrative poetry, and novels– that allows us to examine how lexical fields associated with nation, territory, and alterity are shaped and transformed over almost a century. Treating these materials as a single discursive field makes it possible to relate genres that previous scholarship has tended to analyze in insolation, and to recover shared lexical dynamics that have rarely been considered within a unified framework.

Our approach engages with these contributions but introduces a quantitative perspective based on diachronic lexical analysis and contextual co-occurrence analysis. These procedures enable the identification of larger-scale lexical patterns, contextual proximities, and cross-genre interactions in the symbolic production of the nation.  By combining a multi-genre corpus with computational techniques, we aim to contribute a more integrated view of 19th century Argentine discourse and its evolving representations of national identity. In addition, these methods have not yet been applied to a large-scale corpus of Argentine texts written by Argentine authors, which this project seeks to address.

SPECIFIC OBJECTIVES

The analysis is structured around four recurrent semantic domains that structure 19th discourses on Argentine nationhood: (i) the lexicon of nation and discursive alterity –nación, patria, pueblo, extranjero, indio, and civilización–, which reveals shifting configurations of collective belonging; (ii) space and lexical-semantic territory –campo, ciudad, frontera, país, and pampa– approaching “territory” as a lexical-semantic construct rather than as a material or administrative category; (iii) discursive marginality and difference, analysing the visibility, reconfiguration, and decline of figures as the indio and the gaucho strictly in terms and their lexical and network-based presence, in dialogue with critical traditions on national types (Ortiz 2023, Hanway 2003, Iriarte 2020); and (iv) civilizations and barbarism, examining the resemanticization of barbarie, orden, educación, and progreso between 1830 and 1910.

Together, these domains provide a coherent framework for tracing long-term lexical transformation across genres and historical contexts.

METHODOLOGY AND TEXT COLLECTION

Corpus configuration and editorial criteria

The corpus comprises 61 texts published between 1830 and 1910 and brings together a broad diversity of genres and discursive positions. In its current stage, the dataset contains 2,283,786 word–tokens, distributed across genres such as diaries, memoirs, travelogues, political essays, narrative poetry, novels and short stories (Fig. 1). This range allows us to integrate voices involved both in the representation of the territory –expedition diaries, administrative reports, and frontier memoirs– alongside literary texts that encode urban sensibilities, political tensions, or depictions of social life. Because the corpus is intentionally heterogeneous, its internal composition varies substantially across historical periods, and the results are interpreted as exploratory lexical-semantic tendencies, not as exhaustive representations of 19th century Argentine discourse.

Figure Figure 1. Distribution of corpus size over time based on token counts.

PRELIMINARY RESULTS

The preliminary results indicate that Argentine identity emerges as a discursive process in constant redefinition, articulated around symbolic oppositions that shift as the country aims at modernizing its political, economic and social structures (Fig.2).

Figure Figure 2. Frequencies of selected lexical items across four historical blocks (1830–1910) expressed as occurrences per 10,000 words.

This broader tendency becomes clearer when focusing on the changing collocational environment of nación across historical blocks (Fig.3)

Figure Figure 3. Ego-networks of nación across four historical blocks (1830–1910), based on sentence–level co-occurrences weighted by LogDice

In texts from the 1830–1849 block –dominated by expedition diaries and frontier narratives– associations between nation and a territorial, political, and indigenous-alterity lexicon predominate. Terms such as indio, reducción, salvaje, cacique, or pueblo, together with verbs such as construir, confederar and ocupar, appear closely intertwined in co-ocurrence patterns, shaping an early national imaginary grounded in the contested delimitation of territory, the unstable articulation of the political community, and the construction of the indigenous subject as internal difference.

In the mid-century block, coinciding with the consolidation of civilizational ideology, the lexicon begins to shift, although not in a strictly linear manner. While traces of civilización/barbarie matrix remain visible (argentino, raza, nativo, or civilización), the strongest associations inreasingly point to a juridical and institutional vocabulary. Terms such as tratado, derecho, pacto, administración, justicia, leyreconocer, corresponder, declarar, formar gain prominence, articulate a semantic field linked to the formalizations of the political community and consolidation of the nations as legal and administrative structure. In this sense, the co-occurrence patterns remain compatible with the ideological horizon associated with Sarmiento (1845), but they also suggest a broader process of juridical and political reorganization.

In the later blocks, the strongest collocates point to a historicist and cultural resemanticization of the national imaginary. Words such as raza, tradición, héroe, genio, leyenda, civilización, and evolución become central, alongside markers such as nativo, indígena, and quechua. This shift suggest that the nation is increasingly articulated through symbolic ancestry, collective memory and cultural heritage, reconfiguring earlier frontier and juridical frameworks within a broader discourse of historical identity. By the turn of the century, this repertoire coexists with language of political order, progress, and modernization, indicating a gradual reorientation of the national imaginary toward more institutional and future-oriented forms of cohesion. In the final block (1890–1910), it gives greater prominence to a more explicitly patriotic and civic vocabulary, alongside a future-oriented rhetoric of collective projection.

Whitin this multi-genre corpus, the collocational profiles shown in Fig.4 reinforce this broader transition. The figure of the gaucho –prominent in poetry, narrative, and drama– undergoes a marked reconfiguration, moving from discursive marginality to national emblem, while the indio is increasingly relegated to a residual or archaeological presence, especially in texts from the early 1900s. The co-ocurrence patterns also reveal a change in the representation of territory: campo loses contextual prominence to ciudad, whose associations point to modernizing practices, commerce, cafés, social life, and public space. Together, these patterns indicate a gradual reorientation of the national imaginary from rural territoriality to urban sociability.

Figure Figure 4. Lexical contexts of indio and gaucho by historical blocks. Top 20 co-ocurrences (nouns and adjectives)

CONCLUSIONS

These preliminary results we present as short paper demonstrate the value of digital humanities methods integrating multi-genre corpora into a unified analytical framework and for visualizing discursive transformations from a diachronic perspective. Even though preliminary, the results show that Argentine national identity is constructed through dynamic lexical constellations circulating across heterogeneous genres, and that quantitative techniques reveal relational patterns that complement and refine traditional critical interpretations (Hanway 2003). The distinction between ideological categories and their discursive manifestations proves especially productive for analysing how forms of alterity, marginality, and territorial imaginations are encoded in lexical-semantic patterns.

The conceptual evolution observed does not reflect an abrupt rupture but rather a progressive resemanticization of longstanding ideological axes: the shift from barbarism to order, from the frontier to the state, and from the marginal subject to the modern citizen. The lexical and co-occurrence patterns suggest that Argentine identity is constituted through language, within a complex discursive fabric in which texts do not simply represent the nation but actively participate in shaping its symbolic contours.

The project remains ongoing, with future work aimed at expanding the corpus, refining analytical procedures, and developing comparative perspectives that further illuminate the discursive formation of Argentina between the nineteenth and early twentieth centuries. At a time when the language of the nation is once again circulating forcefully in both institutional discourse and digital media, these historical patterns feel strikingly current. The symbolic oppositions that shaped Argentine identity in the nineteenth century have not disappeared; they still inform present-day debates about who belongs, who is left out, and how national territory is imagined.

References
  1. Calvo Tello, José / Henny-Krahmer, Ulrike / Schöch, Christof (2018): “Textbox. Análisis del léxico mediante corpus literarios”, in: Corbella, Dolores / Fajardo, Alejandro / Langenbacher, Jürgen (eds.): Historia del léxico español y Humanidades Digitales. Bern: Peter Lang 225–253.
  2. Del Rio Riande, Gimena / Henny-Krahmer, Ulrike (eds.) (2025): ArDraCor: Argentinian Drama Corpus https://dracor.org/ar [28.04.2026].
  3. Hanway, Nancy (2003): Embodying Argentina: Body, Space and Nation in 19th Century Narrative. Jefferson, NC: McFarland.
  4. Hengchen, Simon / Ros, Ruben / Marjanen, Jani / Tolonen, Mikko (2021): “A data-driven approach to studying changing vocabularies in historical newspaper collections”, in: Digital Scholarship in the Humanities 36, 2: 329–346.
  5. Henny-Krahmer, Ulrike (2018): “Exploration of sentiments and genre in Spanish American novels”, in: Digital Humanities 2018. Puentes–Bridges https://dh2018.adho.org/exploration-of-sentiments-and-genre-in-spanish-american-novels/ [28.04.2026].
  6. Henny-Krahmer, Ulrike (ed.) (2017a): Collection of 19th Century Spanish-American Novels (1880–1916). CLiGS https://github.com/cligs/textbox/master/spanish/novela-hispanoamericana/ [28.04.2026].
  7. Henny-Krahmer, Ulrike (ed.) (2017b): Bib-ACMé: Bibliografía digital de novelas argentinas, cubanas y mexicanas (1830–1910). CLiGS http://bibacme.cligs.digital-humanities.de/ [28.04.2026].
  8. Henny-Krahmer, Ulrike / Neuber, Frederike (2017): Criteria for Reviewing Digital Text Collections (Version 1.0). Köln: Institut für Dokumentologie und Editorik (IDE) https://www.i-d-e.de/publikationen/weitereschriften/criteria-text-collections-version-1-0/ [28.04.2026].
  9. Jänicke, Stefan (2016): “Visual text analysis in digital humanities”, in: Computer Graphics Forum 35, 3: 509–525. DOI: 10.1111/cgf.12873.
  10. Jänicke, Stefan / Franzini, Greta / Cheema, Muhammad Faisal / Scheuermann, Gerik (2015): “On close and distant reading in digital humanities: A survey and future challenges”, in: Computer Graphics Forum 34, 3: 201–221. DOI: 10.1111/cgf.12620.
  11. Jockers, Matthew L. (2013): Macroanalysis: Digital Methods and Literary History. Urbana / Chicago / Springfield: University of Illinois Press http://www.jstor.org/stable/10.5406/j.ctt2jcc3m [28.04.2026].
  12. Martínez Estrada, Ezequiel (1933): Radiografía de La Pampa. Buenos Aires: Babel.
  13. Martínez Estrada, Ezequiel (1940): La cabeza de Goliat. Buenos Aires: Club del Libro Amigos del Libro Americano.
  14. Ortale, María Celia (2022): “Desplazamientos discursivos en la recopilación de las Obras completas de José Hernández: ¿La desaparición del héroe?”, in: Orbis Tertius 27, 35: 1–17 https://www.memoria.fahce.unlp.edu.ar/art_revistas/pr.16070/pr.16070.pdf [28.04.2026].
  15. Ortiz, Claudia M. (2023): “Mapa de actores para la lectura de cartas e informes: La Expedición Lista (1886–1887)”, in: Estudios históricos 17, 34: 139–162 https://www.redalyc.org/articulo.oa?id=379977805009 [28.04.2026].
  16. Padró, Lluís / Stanilovsky, Evgeny (2012): “FreeLing 3.0: Towards Wider Multilinguality”, in: Calzolari, Nicoletta / Choukri, Khalid / Declerck, Thierry / Dogan, Mehmet Uğur / Maegaard, Bente / Mariani, Joseph / Odijk, Jan / Piperidis, Stelios (eds.): Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC 2012). Istanbul: European Language Resources Association 2473–2479Rojo, Guillermo (2021): Introducción a la lingüística de corpus en español. London / New York: Routledge.
  17. Rychlý, Pavel (2008): “A lexicographer-friendly association score”, in: Sojka, Petr / Horák, Aleš (eds.): Proceedings of RASLAN 2008: Recent Advances in Slavonic Natural Language Processing. Brno: Masaryk University 6–9 https://www.sketchengine.eu/wp-content/uploads/2015/03/Lexicographer-Friendly_2008.pdf [28.04.2026].
  18. Sarmiento, Domingo Faustino (1845): Civilización i barbarie. Vida de Juan Facundo Quiroga, i aspecto físico, costumbres i hábitos de la República Argentina. Santiago de Chile: Imprenta del Progreso.
  19. Schöch, Christof (2017): “Quantitative Analyse”, in: Jannidis, Fotis / Kohle, Hubertus / Rehbein, Malte (eds.): Digital Humanities: Eine Einführung. Stuttgart: Metzler 279–298.
  20. Schöch, Christof / Henny-Krahmer, Ulrike / Calvo Tello, José / Popp, Sandra (2016): “Topic, genre, text: Topics im Textverlauf von Untergattungen des spanischen und hispanoamerikanischen Romans (1880–1930)”, in: DHd 2016. Leipzig: Universität Leipzig 235–239 http://www.dhd2016.de/abstracts/vortr%C3%A4ge-055.html [28.04.2026].