DH 2026

Daejeon, July 27–31

Wed, July 2916:30–18:00S048106
Short Paper

From Princeton Prosody Archive to Reusable Infrastructure: A Content-Agnostic, Provider-Agnostic Platform for Full-Text Searchable Collections

Hao Tan
Center for Digital Humanities, Princeton University, United States of America · tanhao@princeton.edu
Rebecca Koeser
Center for Digital Humanities, Princeton University, United States of America · rebecca.s.koeser@princeton.edu
Mary Naydan
Center for Digital Humanities, Princeton University, United States of America · mnaydan@princeton.edu
Meredith Martin
Center for Digital Humanities, Princeton University, United States of America · mm4@princeton.edu

Digital humanities projects frequently converge on the same practical goal—making digitized cultural materials searchable, interpretable, and engaging for communities—yet they often arrive there through bespoke platforms shaped by a single corpus, genre, institution, or ingestion pipeline. Even when source code is shared, reuse commonly fails in practice because “what looks like code” is only the visible layer; the hidden layer includes decisions and assumptions about providers, storage layouts, authentication, metadata norms, interface language, and editorial workflows. The result is repeated reinvention: similar search interfaces, similar ingestion scripts, similar curation mechanisms, built again and again across projects that could otherwise share infrastructure. This tension is directly relevant to DH2026’s theme of “Engagement” because the ability to engage broader communities is constrained not only by interface design, but by whether under-resourced projects can realistically adopt a platform at all (Risam and Gil 2022; Maron and Pickle 2014; Martin 2025).

This short paper presents PPA Reuse, an ongoing effort that generalizes the software stack behind the Princeton Prosody Archive (PPA) into a genre-agnostic and provider-agnostic toolkit for building full-text searchable archives across heterogeneous corpora. Early design work expands beyond PPA’s current sources to support additional public repositories (e.g., Project Gutenberg and the Internet Archive). PPA is a public archive on historical debates about prosody, pedagogy, and phonetics, drawing on thousands of digitized books from multiple providers, including HathiTrust, Gale, and EEBO-TCP. Related discussions have framed PPA in terms of working around “walled gardens” and treating archive-building as an end-to-end workflow rather than a single interface (Naydan, Koeser, and Martin 2024). Prior work describes PPA’s scholarly motivations as well as technical challenges around legacy metadata and “typographically unique” sources that resist clean OCR (Martin et al. 2018). PPA Reuse treats this history not as a domain constraint but as an engineering opportunity. It asks which parts of a mature DH platform can be abstracted into reusable components without flattening the scholarly and curatorial affordances that made the original project valuable.

The project is driven by a close reading of what has made previous generalization efforts fall short. Platforms such as Omeka have lowered barriers for specific use cases, but each embeds assumptions—about metadata schemas, hosting environments, or ingestion pipelines—that become liabilities when a new project diverges even modestly from the template. The Endings Project has likewise foregrounded sustainability constraints that push toward static output, a legitimate trade-off that nonetheless limits dynamic search and faceting (Holmes, Jenstad, and Huculak 2023). Against this backdrop, PPA Reuse identifies three kinds of variation that are normal rather than exceptional in humanities data, and for which existing platforms provide no systematic account. The first is source variation: corpora are built from many different providers and encodings (institutional digitization pipelines, public repositories, TEI editions, bibliographic exports, OCR packages), and researchers often need to mix them. This variation is not only structural but temporal: provider-side updates can introduce substantial drift in what a corpus contains over time, complicating reproducibility and follow-up research (Koeser et al. 2025). The second is discovery variation: different genres demand different ways of searching and browsing. A cookbook collection might need robust handling of tables, measures, and ingredient-like facets; fiction might emphasize editions, paratext, serialization, and narrative navigation; newspapers may foreground date, place, and typographic variants. The third is interpretive variation: in many culturally important materials, the page’s spatial organization carries meaning and supports claims, as tables separate categories, marginalia mark commentary and authority, and notation or diagrams encode semantics through position and alignment, so treating page form simply as “OCR noise” erases what scholars can legitimately cite and interpret. What distinguishes this framing from prior platform critiques is the claim that these three variations are structurally coupled: a platform that solves source variation alone, while hard-coding discovery and interpretive assumptions, still forces reinvention at the layers that matter most for scholarship.

Accordingly, PPA Reuse’s development focuses on reframing which aspects of an archive are configurable and reviewable. The project is building toward (a) clearer boundaries between “where the data comes from” and “how the archive behaves,” so other scholars can adapt the PPA infrastructure for their own corpora without rewriting core assumptions, and (b) mechanisms that let projects express their interpretive priorities—what should be searchable, what should be faceted, what should be highlighted as form-significant, what should be restricted or contextualized—as explicit configuration and curatorial decisions rather than hard-coded template logic. This orientation aligns with minimal-computing arguments that infrastructure should lower barriers to participation and reduce the hidden labor required to keep projects running over time (Risam and Gil 2022), while also acknowledging that engagement includes stewardship: making provenance and curatorial choices visible rather than silently embedding them in software defaults.

PPA Reuse treats community-facing participation as part of the platform’s long-term direction, but it does not assume a single engagement model. Instead, it aims to support a range of editorial and community practices, including narrative pages and curated groupings for teaching and public scholarship, clear provenance statements for reuse and accountability, and pathways for community commentary where appropriate, informed by governance and care principles (Carroll et al. 2020). It designs these capabilities as optional, feature-flagged components—a design choice that directly lowers the cost of adoption for under-resourced or small-corpus projects, which cannot maintain features they do not need. Alongside this modularity, the project has also prioritized making the platform itself easier to deploy and run, reducing the operational overhead that has historically made even well-documented DH infrastructure inaccessible to smaller teams. This modularity also raises a question the paper aims to bring to the conference: what criteria should determine whether a platform is ready to be adopted as a reusable model? Prior platforms have become de facto standards through promotion, institutional adoption, or community recognition rather than through explicit fitness criteria. PPA Reuse proposes that reusability should be evaluated against the three variations described above, but we invite the community to contest or extend that standard.

This short paper contributes a reusable-infrastructure report on early software design work intended to open discussion in the short-paper session. We will use the presentation to (1) critically map the recurring failure points that prevent DH platforms from being reused across genres, with reference to existing platforms, (2) outline design principles for “content-agnostic” and “provider-agnostic” archive infrastructure, and (3) share small pilot scenarios (cookbook, fiction, and notation-heavy subsets) as concrete thought experiments for what configuration and evidence preservation must support. We invite feedback on what DH communities should consider meaningful evidence that a platform is truly reusable: which variations matter most, which curatorial decisions must remain visible, and which governance practices should be supported as built-in platform capabilities and defaults, rather than left to ad hoc project policy.

In the spirit of DH2026’s theme of Engagement, we argue that a platform truly “engages” when more projects, including under-resourced teams and small-corpus initiatives, can adopt it, and when genre-specific evidence, such as tables, recipes, marginalia, and notation, remains legible rather than flattened into text. Presenting PPA Reuse at an early stage is therefore an invitation to collective critique on what engagement should require from reusable archive infrastructure, before those requirements harden into defaults.

References
  1. Carroll, Stephanie Russo et al. (2020): "The CARE Principles for Indigenous Data Governance", in: Data Science Journal 19, 43.
  2. Holmes, Martin / Jenstad, Janelle / Huculak, J. Matthew (2023): "Introduction to Special Issue: Project Resiliency in the Digital Humanities", in: Digital Humanities Quarterly 17, 1.
  3. Koeser, Rebecca Sutton / Budak, Nick / Hicks, Benjamin W. / Doroudian, Gissoo (2018): "Princeton-CDH/ppa-django: v3.0 Initial Public Release". Zenodo. DOI: 10.5281/zenodo.2400705.
  4. Koeser, Rebecca Sutton / Budak, Nick / Heuser, Ryan / Thompson, Laure / Tan, Hao / Hicks, Meg / Doroudian, Gissoo / Bansal, Vineet / McElwee, Kevin (2026): "ppa-django". Zenodo. DOI: 10.5281/zenodo.18174442.
  5. Koeser, Rebecca Sutton / Naydan, Mary / Martin, Meredith (2025): "Unstable Data and the Unusual Case of the Prosody Excerpt in the Digital Library", in: Arnold, Taylor / Fantoli, Margherita / Ros, Ruben (eds.): Computational Humanities Research 2025. Anthology of Computers and the Humanities, vol. 3: 1390–1403. DOI: 10.63744/cTWwVSItf41f.
  6. Maron, Nancy L. / Pickle, Sarah (2014): Sustaining the Digital Humanities: Host Institution Support beyond the Start-Up Phase. Ithaka S+R.
  7. Martin, Meredith (2025): Poetry's Data. Princeton: Princeton University Press.
  8. Martin, Meredith / Wilson, Meagan / Naydan, Mary (2018): "Princeton Prosody Archive: Rebuilding the Collection and User Interface". Poster presented at DH2018, Mexico City, June 27, 2018.
  9. Naydan, Mary / Koeser, Rebecca Sutton / Martin, Meredith (2024): "Beyond the Walled Gardens: Reinventing the Digital Research Landscape with the Princeton Prosody Archive". Poster presented at DH2024, Washington, D.C., August 9, 2024.
  10. Princeton Prosody Archive (2018): Version 3.16.0. Center for Digital Humanities at Princeton http://prosody.princeton.edu [08.05.2026].
  11. Risam, Roopika / Gil, Alex (2022): "Introduction: The Questions of Minimal Computing", in: Digital Humanities Quarterly 16, 2.