DH 2026

Daejeon, July 27–31

Thu, July 3016:30–18:00S067106
Short Paper

What the Archive Remembers: Geography, Race, and Language in Chronicling America

David Bishop
University of Illinois Urbana Champaign, United States of America · Davidrb3@illinois.edu

This paper examines how large-scale digital newspaper archives shape scholarly engagement with the literary past, arguing that such archives function as powerful—but uneven—systems of cultural memory. Projects such as the Library of Congress’s Chronicling America have expanded access to historical newspapers and enabled new forms of computational research in literary history and print culture. Yet they also reconfigure what counts as “the archive” through selective digitization, metadata design, OCR legibility, and infrastructural constraints. These systems do not merely preserve memory; they actively produce it. Understanding how this production occurs is essential for responsible engagement with computational methods, particularly as AI-assisted tools increasingly rely on large archival corpora.

This project advances current scholarship that interrogates the digital archive as a historically contingent, methodologically constrained system rather than a neutral repository of evidence. It sits between the broad computational and editorial reflections of Cordell et al.’s “Editing a Paper,” which conceptualize digitized newspaper collections as editorially constructed systems across four large-scale corpora—Chronicling America, Gale Nineteenth Century U.S. Newspapers, ProQuest American Periodicals Series, and the Internet Archive—and Cordell’s more granular analysis in “Q i-jtb ‘The Raven,’” which examines how the National Digital Newspaper Program’s (NDNP) grant politics at the state level—particularly in Pennsylvania—condition patterns of digitization and literary visibility. Rather than treating digitized newspapers as transparent evidence, this project approaches them as historically produced and methodologically constrained archives whose structures shape the interpretive claims that can be made from them.

This approach situates the project within a broader body of digital humanities scholarship concerned with the construction and limits of digital archives. Emily Bell and M. H. Beals observe that beneath the polished interfaces of digitized newspaper collections lie decades of curatorial decision-making that determine what users can search for—and what they will find. Benjamin Fagan has demonstrated how racially uneven digitization practices have produced distorted digital archives, while Katherine Bode argues that datasets should be treated as interpretive texts rather than transparent evidence. Taken together, this work reframes digitized archives as constructed systems whose biases must be made legible before they can be analytically productive.

Focusing on Chronicling America as a case study, this paper reports on an ongoing analysis drawing on bulk newspaper metadata made publicly available by the Library of Congress for both Chronicling America and the U.S. Newspaper Directory. While the USND records over 119,000 known newspaper titles published since 1690, Chronicling America includes just over 4,000 digitized titles. Rather than treating this disparity as a methodological problem to be solved, I read the digital archive as an interpretive artifact whose gaps and distortions shape what forms of literary history become visible at scale. This project employs computational aggregation and advances a reusable analytical workflow grounded in patterns of coverage, visibility, and absence.

To make this “memory work” legible, the project develops a three-pronged audit of representational bias across geography, race, and language. The goal is not to correct the archive into representativeness, but to provide a framework for critically engaging its structure before drawing conclusions from large-scale computational analysis.

The geographic analysis reveals a striking inversion of historical print geographies. Whereas newspaper production between 1789 and 1963 was concentrated in metropolitan centers such as New York, Illinois, and Pennsylvania, the digitized archive disproportionately foregrounds southern and interior states. These patterns reflect not historical output alone, but uneven digitization infrastructures, funding pathways, and preservation priorities. Digitization thus operates as a second wave of archival selection, reshaping national cultural memory according to contemporary institutional conditions.

The analysis of the Black press demonstrates why statistical measures of representativeness must be interpreted with care. Fagan’s critique of the “racial politics of periodical digitization” showed that, as late as 2016, Chronicling America made no Black newspapers available for the period prior to 1865—an absence with profound consequences for DH projects working at scale. While my findings show that the numerical presence of Black newspapers has since increased substantially, numerical overrepresentation should not be conflated with humanistic representation.

The language analysis shows that digitization also remaps the United States’ multilingual cultural record. While the USND catalogs newspaper titles across a broad language ecology, Chronicling America collapses this multilinguistic landscape. German, Spanish, and French appear amplified, while many languages—including Chinese, Korean, Greek, Ukrainian, Portuguese, Croatian, Dutch, Armenian, and others—disappear entirely from the digitized corpus. These absences delimit what kinds of cultural memory become available for computational research and risk being mistaken for historical silence rather than infrastructural absence.

Beyond diagnosing distortion, I propose a methodological intervention for working with large-scale newspaper archives that treats representational bias as an analytic signal rather than a limitation to be corrected or bracketed. I argue for reading Chronicling America as a springboard for inquiry by operationalizing patterns of coverage, visibility, and absence as inputs to research design. Within this framework, state-level variation becomes analytically meaningful: inconsistent digitization—shaped by state–federal partnerships, regional consortia, and local preservation infrastructures—is read as evidence of where specific print networks are rendered computationally legible and where others remain structurally obscured. This approach is especially productive for research on the Black press. In the course of this research, I identified and recovered the works of Augustus M. Hodges in preparation for a digital critical edition. Hodges was a prolific Black writer whose serialized fiction appeared frequently in The Indianapolis Freeman. His visibility in the digital archive reflects archival contingency rather than representativeness and functions as a trigger for targeted archival investigation. The language audit further reveals how digitization initiatives encode institutional priorities about cultural diversity, signaling which traditions are rendered machine-readable and which remain dependent on alternative archival strategies for recovery.

This project theorizes Chronicling America as a system of remembering, foregrounding questions of curation, ownership, and access by asking whose memories are preserved, which forms of cultural expression become machine-readable, and for whose benefit. Ultimately, it demonstrates how engagement with the past is structured—sometimes enabled, sometimes foreclosed—by the archival and computational systems that mediate scholarly and algorithmic interpretation.

References
  1. Beals, M. H. / Emily Bell. (2020): Introducing The Atlas of Digitised Newspapers and Metadata, in RSVP. rs4vp.org/introducing-the-atlas-of-digitised-newspapers-and-metadata/.
  2. Bode, Katherine. (2018): A World of Fiction: Digital Collections and the Future of Literary History. University of Michigan Press.
  3. Chronicling America: Historic American Newspapers. Library of Congress. OCR text dataset via Library of Congress API. https://chroniclingamerica.loc.gov/.
  4. Cordell, Ryan. (2017): “Q I-jtb the Raven: Taking Dirty OCR Seriously.” in Book History, vol. 20, no. 1, pp. 188–225, https://doi.org/10.1353/bh.2017.0006.
  5. Cordell, Ryan, et al. (2023): "Editing a Paper." Going the Rounds: Virality in Nineteenth-Century American Newspapers, University of Minnesota Press. manifold.umn.edu/read/editing-a-paper/section/fc0597a3-5fe1-439c-86f8-0e47e8a55208. Accessed 9 Feb. 2025.
  6. Fagan, Benjamin. (2016): “Chronicling White America.” In American Periodicals, vol. 26, no. 1, pp. 10–13. EBSCOhost, doi.org/10.2307/44630659.
  7. Library of Congress. U.S. Newspaper Directory, 1690–Present. Library of Congress, https://www.loc.gov/newspapers/.