Introduction Anthologies serve both as pedagogical tools and as repositories of cultural memory. While their selections may initially mirror an editor’s personal taste, their lasting impact is to define a poet’s image over time. However, this selection process inherently creates a representational gap; by reducing an entire career to a few representative works, editors set interpretive expectations that might not fully align with a poet's complete body of work. This study shifts from the idea that anthologies may be biased to the need to quantify that bias, measuring the structural differences between highly visible selections and the broader poetic archive.
Data & Corpus The study analyzes the textual patterns of several 20th-century Spanish poets across two distinct datasets. The first dataset includes their complete poetic works, digitized and curated to cover their entire published output. The second dataset features selections from influential sources, centered on the key volume edited by Gustavo Correa (1980). By comparing the comprehensive baseline of the complete oeuvres with these curated selections, we investigate how editorial mechanisms bring certain poems into sustained visibility while leaving the majority of the archive underexamined.
Methodology Methodologically, the study uses a multi-layered computational approach to understand editorial intervention. We combine large-scale multivariate analysis, including Principal Component Analysis (PCA) using the R stylo package (Eder et al., 2016), with detailed feature extraction. This includes Zeta analysis to identify distinctive authorial tokens and recent LLM-based methodologies (Martínez et al., 2025) to measure cognitive and affective dimensions—such as concreteness and valence—across lexical and emotional levels. This combined method enables us to measure structural differences and identify which aspects of the poets' voices are favored or overlooked in curated collections.
Initial Findings Initial findings from a pilot analysis reveal a clear structural difference. PCA shows that the vector distances between the poets’ complete works are relatively close, indicating a shared literary foundation or generational connection. Conversely, the anthology selections show much greater spread in the vector space, suggesting that they tend to include highly distinctive poems that emphasize stylistic differences and create a more defined, yet artificial, individuality for each poet. Significant variances in emotional tone and imagery empirically highlight the influence of editorial subjectivity in shaping these representative profiles. The specific patterns of these stylistic distortions will be detailed in the presentation.
Conclusion By systematically comparing editorially chosen poems with the larger poetic archive, this study advances ongoing debates in digital humanities and literary historiography. It seeks to establish a multidimensional and empirically grounded framework for understanding the complex relationships between editorial curation, the forces behind literary visibility, and the vast, silent body of unread works. We open up the possibility to re-read a more complete and diverse literary history.
References
Correa, Gustavo (ed.) (1980): Antología de la poesía española (1900-1980). Madrid: Editorial Gredos.
Eder, Maciej / Rybicki, Jan / Kestemont, Mike (2016): “Stylometry with R: a package for computational text analysis”, in: The R Journal 8, 1: 107–121.
Martínez, Gabriel / Conde, Juan / Reviriego, Pedro / Brysbaert, Marc (2025): “AI-generated estimates of familiarity, concreteness, valence, and arousal for over 100,000 Spanish words”, in: Quarterly Journal of Experimental Psychology 78, 10: 2272–2283. DOI: 10.1177/17470218241251336.