Daejeon, July 27–31
Known as "avisos de fuga," runaway advertisements for domestic servants constitute one of the richest repositories of personal biometric information in post-Independence Latin America. This presentation examines the early formation of physical identification categories that emerged from descriptions in approximately 500+ unique runaway advertisements published in Lima's El Comercio newspaper between 1839 and 1868. Using interdisciplinary methods from Data Science, Science and Technology Studies (STS), Social History, and Digital Humanities (including Natural Language Processing), we analyze how these advertisements became platforms for constructing personal data at a quotidian level within print media. We argue that runaway advertisements served as crucial precursors to modern biometric systems, establishing descriptive categories that would later be systematized into digital databases and algorithmic identification systems. Their intersectional descriptions of mestizo-Indigenous domestic workers reveal how racial, class, and gender biases became embedded within identification practices through the repetitive use of physical descriptors. In nineteenth-century Perú, "cholo" referred to individuals of mixed Spanish and Indigenous descent, a term that emerged from colonial caste systems and carried deeply discriminatory connotations associated with racial and class hierarchies.
Hence, in this work, we provide preliminary evidence showing how analog systems of personal identification in postcolonial contexts generated the categorical frameworks that continue to shape contemporary biometric technologies, often reproducing historical prejudices against vulnerable communities through seemingly neutral technical processes. This rare historical analysis of pre-digital biometric category formation in Latin America gives us critical insights for addressing algorithmic bias and technological inequalities in the Global South.
Our interdisciplinary approach builds on critical scholarship examining the intersection of race, technology, and surveillance. Simone Browne's Dark Matters (2015) and Ruha Benjamin's Race After Technology (2019) demonstrate how surveillance systems have historically targeted marginalized bodies, a pattern we trace back to nineteenth-century print media. Safiya Umoja Noble's Algorithms of Oppression (2018) reveals how contemporary search engines perpetuate racial bias: a phenomenon whose genealogy we would locate in the categorical frameworks established by runaway advertisements. Our research extends these critical interventions by providing historical depth to conversations about algorithmic justice, showing how biometric categorization emerged from unequal power relations long before the digital age.
The project employs computational text analysis methods that bridge traditional historical research with Digital Humanities approaches. Following Melanie Walsh's framework in Introduction to Cultural Analytics & Python (2020), we apply NLP techniques such as TF-IDF (Term Frequency-Inverse Document Frequency) to identify the most significant descriptive terms across the corpus and Topic Modeling to trace thematic patterns and their temporal evolution. Walsh's emphasis on making computational methods accessible to humanities scholars informs our methodological transparency and our commitment to creating reproducible research workflows. This approach allows us to move between "distant reading" of large textual corpora and "close reading" of individual advertisements, revealing both systemic patterns and particular human stories (Moretti, 2013).
A central contribution of this project is the creation of a publicly available database in CSV format containing structured information from all 500+ runaway advertisements. This dataset represents the first systematic digital collection of biometric descriptions from nineteenth-century Latin America, filling a significant gap in both Digital Humanities resources and Latin American historical data. The database includes fields for physical descriptors, occupational information, ethnic/racial categorizations, geographic references, and temporal markers, all carefully extracted through our combined computational and qualitative analysis. By making this resource openly accessible, we aim to enable future research on topics ranging from historical demography to computational linguistics, from labor history to critical race studies.
As Catherine D’Ignazio and Lauren Klein argue in Data Feminism (2020), making visible the origins and construction of data systems is crucial for challenging their naturalization and exposing inequalities. Our database documentation will include detailed metadata about extraction processes, coding decisions, and interpretive choices, modeling transparent data practices that Digital Humanities can bring to historical research. By revealing the constructed nature of identification categories in our historical corpus, we hope to contribute to ongoing efforts toward more equitable data systems. Through this short paper, we aim to demonstrate how Digital Humanities methodologies offer powerful tools for studying historical texts in new ways, revealing patterns invisible to traditional archival research alone.
Our preliminary analysis suggests that the avisos de fuga show distinct temporal shifts in descriptive vocabulary, reflecting changing racial ideologies and identification practices across the long nineteenth century. We expect to identify core sets of phenotypic descriptors that appear with striking consistency, suggesting the emergence of standardized categorical frameworks even in this ostensibly informal context. Furthermore, we anticipate finding intersectional patterns where descriptions varied significantly based on the presumed ethnic/racial identity of the fugitive, revealing how power asymmetries shaped information capture from its earliest moments.
This research contributes to ongoing conversations about technological justice, demonstrating that contemporary concerns about algorithmic bias have deep roots in colonial and postcolonial identification practices, building on the work of Data Justice (Dencik et al., 2022), where the problem of historically situated data (in)justice is presented as a continuation of previous forms of oppression. Our work shows how Digital Humanities can illuminate the genealogies of present-day inequalities, providing essential context for those working to create more equitable technological futures, as well as providing valuable insights to gather further understanding about the historical roots of discrimination in the Global South.