Daejeon, July 27–31
At a recent lecture on “Better Than Best Practices in Data Work for Culture and Society,” Alex Gil urged scholars to rethink how we manage, analyze, and describe historical cultural data—especially records shaped by the colonial language of the nineteenth century. He challenged researchers to move beyond inherited schemas such as Dublin Core and instead create controlled vocabularies that serve descendant communities. Roopika Risam similarly warns that “epistemic violence of colonialism … is implicated in colonial forms of knowledge production” (51), a concern reflected in the linked data structures of platforms such as OCLC WorldCat (Metz 2025). Recent approaches—including “kinfolkology” (Kinfolkology), naming practices (Hawthorne et al. 2025), and Feminist Bibliography’s structures of care (Ozment 2025)—underline the need to treat historical records as traces of lived experience rather than neutral data. Metadata shapes discoverability, but it also shapes the corpora used for computational analysis that can challenge long-standing assumptions about the nineteenth-century British canon.
This short paper introduces a modest dataset and corpus built from widely circulated nineteenth-century literary annuals—serial publications that helped define and export British exceptionalism. Although nearly every canonical author wrote for these annuals, their literary and visual materials remain scattered; few libraries hold complete runs. Published from 1823 to 1847, these annuals illuminate binding practices, editorial decisions, and the economics of canon formation. They also circulated racist and Orientalist depictions of British colonial subjects, including those in annuals printed in Calcutta by an East India Company officer. Their poetry, essays, dramatic pieces, and short stories—paired with steel-plate engravings—attempted to shape readers’ perceptions of both domestic and colonial spaces. Because the genre is so visually and textually hybrid, interpretive work on the annuals requires access not only to the literary content but also to the engravings that reinforced their ideological framing.
To address long-standing accessibility and discoverability challenges, this project collaborated with the Internet Archive to digitize a complete private collection of literary annuals. Working through an established nonprofit with existing linked-data infrastructure enabled additional critical cataloging interventions. The project corrected earlier metadata, highlighted colonial perspectives in translations and literary representations of enslaved communities, and created an OCR-based corpus. This involved the preparation of MarcXML records for each volume, ensuring that authors, engravers, and artists previously omitted or inconsistently attributed could be located through linked data. Because this dataset intentionally remains small—fewer than one hundred volumes—it affords a level of manual care and contextual correction that would be impossible at scale. The size of the corpus also aligns with feminist and postcolonial commitments to descriptive precision: each item can be interpreted, corrected, and contextualized on its own terms.
A new workflow supported tagging of 1,000 steel-plate engravings with controlled vocabulary that specifies, for example, “Indian woman working” or “Bengali boy reading,” linked to the accompanying literary text. Many engravings required not only subject classification but also acknowledgment of colonial framing—for instance, noting when an “Indian woman working” is depicted in a stylized or stereotyped position that does not reflect lived labor practices. To accomplish this, the project used a quadrant-based machine-learning process: each engraving was divided into four to six segments so that localized features could be described precisely when training data was limited. The goal was not automation but improved interpretive accuracy in a small-data context. These methods work toward dismantling “the colonial dynamics of the digital cultural record [and] produce new ways of knowing” (Risam 57). The workflow also enables scholars to pair specific images with their corresponding texts, restoring the original visual-literary relationships that shaped nineteenth-century reading practices.
As I prepared MarcXML files for Internet Archive crosswalks, a limitation became clear: OCLC records will not update, even when errors are corrected or when colonial descriptors are replaced with respectful alternatives. However, subject-area indexes maintained by professional organizations will surface the enhanced metadata, and uploading MarcXML files to Wikidata will support automated ingestion. This ensures that authors, engravers, and artists associated with literary annuals become more discoverable—an important step at a moment shaped by generative-AI distortions, which often reproduce the biases of incomplete or incorrect datasets. Linked open data allows these materials to be read alongside more canonical nineteenth-century texts, expanding the scope of what can be computationally analyzed without reinscribing harmful descriptive practices.
Though neither a large-scale dataset nor a traditional digital project, this collaboration demonstrates how linked open data can support new interpretive work across 84,900 pages of text and 1,000 engravings while resisting the reproduction of colonialist practices. As a small-data project grounded in critical cataloging, it offers a model for how historical serial publications can be responsibly digitized, enriched, and mobilized for both humanistic interpretation and computational use.