DH 2026

Daejeon, July 27–31

Poster

EngPLURIBUS: A Multigenre, English Dataset of Book Translations from Multiple Languages

Sarah Griebel
University of Illinois Urbana-Champaign, United States of America · sarahg8@illinois.edu
Glen Layne-Worthey
University of Illinois Urbana-Champaign, United States of America · gworthey@illinois.edu
J. Stephen Downie
University of Illinois Urbana-Champaign, United States of America · jdownie@illinois.edu

Introduction:

Translation studies have a long history in the digital humanities, and the official DH2026 sub-theme “Translating | Translinguality in the Era of AI” is evidence that this continues to be the case. Despite the ongoing interest, there are still gaps in the availability of training and reference corpora needed for research in this area.  Our work seeks to join other important translation datasets in filling those gaps.

This project introduces a new corpus, EngPLURIBUS

Dataset and code for this project can be found at https://github.com/griebels/identifying-translations.

: English-language translations published prior to 1930, drawn from the HathiTrust Digital Library. The corpus includes 19,430 translations from 85 languages, including a subset of manually verified bridge translations. The corpus is constructed along three axes: (1) historical reach, enabling computational translation research into the eighteenth and nineteenth centuries; (2) depth of genre, including both fiction and nonfiction texts; and (3) bibliographic completeness, cross-referencing multiple MARC21 (library catalog metadata) fields to computationally identify translations.

Our work focuses on pre-1930 translation culture, characterized by a shift in theoretical norms from situating the translation closer to the reader (“domestication”) towards bringing the reader closer to the original author (“foreignization”) (Schleiermacher 1813). We also introduce the important concept of “bridge translation” into computational translation discourse.

Related Literature:

Several recent datasets have laid the groundwork for this kind of work. The TRANSCOMP dataset (Erlin et al. 2022) provides over 10,000 translations into English alongside a parallel collection of English-language originals from 1950-2000, derived from the NovelTM dataset (Underwood et al. 2020). While TRANSCOMP focuses on contemporary works, our corpus extends the lineage backward into previous centuries and expands the genre categories.

On another scale, MultiHATHI (Hamilton and Piper 2023) provides a complete multilingual inventory of prose fiction in HathiTrust, covering over 500 languages and exposing asymmetries in global textual availability. While MultiHATHI focuses on the presence (or absence) of certain languages, our project focuses on how texts are connected through language: identifying which texts explicitly circulate across languages and how.

A central challenge in historical corpus construction is that translations are not consistently labeled in bibliographic metadata. Sirin et al. (2024) confront this issue directly in their work on structured multilinguality in HathiTrust, showing that language and translation information frequently appear in non-standardized MARC21 Notes fields rather than in controlled vocabularies. Our research has also shown that even standardized MARC21 fields related to translations are used inconsistently and idiosyncratically in the HathiTrust catalog.

This perspective reframes metadata not as a static truth source but as a noisy and incomplete substrate requiring computational interpretation. Research in metadata enhancement (Karlinska et al. 2024) extends this logic, demonstrating how bibliographic data becomes significantly more powerful when enriched, linked, and restructured for computational reuse.

Clear and structured corpus construction allows further computational work on studying translations, including intertextuality (McGovern et al. 2024), translationese (Ni et al. 2022), equivalence (Porter et al. 2022), and the effects of cultural mediation (Erlin et al. 2025).

Methods:

To construct our corpus, we identified all records in the HathiTrust Digital Library containing or linked to a work containing the MARC21 765, which generally identifies a text as a translation of another work.

Building on the methods developed in TRANSCOMP, we apply regular-expression searches through metadata fields, including title pages, edition statements, notes, and added entry fields, to identify candidate translations.

We restrict the corpus to pre-1930s texts and apply a language-detection algorithm to samples of the Extracted Features Dataset 2.5 (Walsh et al. 2025) to automatically identify language, retaining only texts receiving greater than 0.75 English probability. To determine source and bridge languages, we combine manual and automated approaches. Records indicating multiple source languages are first flagged as potential bridge translations for manual review. For those confirmed as bridge translations, source and intermediary languages are then identified through close reading of the relevant metadata fields. For all others, we apply a scaffolded pipeline that maps MARC21 language codes to ISO 639-2 codes and supplements these with regular-expression searches for explicit language references (e.g., “translated from the [language]”).

Preliminary Conclusions:

Bridge translations are more reliably identifiable in the notes fields than in structured MARC21 relationship fields. While formal linking fields such as Field 765 typically encode only direct source-target relationships, notes frequently contain additional contextual detail, including references to

intermediate languages. Unstructured descriptions offer richer evidence for the detection of indirect translation pathways that are not formally represented elsewhere in the metadata.

Additionally, European languages dominate as source languages, with French and German alone comprising nearly 40% of the total source languages (see Fig.1). It remains unclear whether this pattern reflects historical translation practices of the period or, in part, collection biases within the HathiTrust Digital Library. French emerges as the most prominent bridge language, accounting for over half of all identified bridge translations (see Fig. 2 and Fig. 3), underscoring its central role as a mediating language in historical literary circulation.

Figure 1. Top 15 source languages in the entire corpus, with individual counts.

Figure 2. Bridge translation network graph. Arrows indicate both direction and volume.

Fig 3. Most prevalent bridge and source languages, along with their counts, for the 51 identified bridge translations.

References
  1. Erlin, Matt / Knox, Douglas / Carroll, Claudia / Sushil, Jey / Ussiri, Tumaini / Watanabe, Sadahisa (2025): “Geotropes: Situating Postcolonial Bestsellers in the Global Literary Marketplace”, in: Journal of Cultural Analytics 10, 2. https://doi.org/10.22148/001c.142973.
  2. Erlin, Matt / Piper, Andrew / Knox, Douglas / Pentecost, Stephen / Blank, Allie (2022): “The TRANSCOMP Dataset of Literary Translations from 120 Languages and a Parallel Collection of English-Language Originals”, in: Journal of Open Humanities Data 8: 29. https://doi.org/10.5334/johd.94.
  3. Hamilton, Sil / Piper, Andrew (2023): “MultiHATHI: A Complete Collection of Multilingual Prose Fiction in the HathiTrust Digital Library”, in: Journal of Open Humanities Data. https://doi.org/10.5334/johd.95.
  4. Karlinska, Agnieszka / Rosiński, Cezary / Kubis, Marek / Hubar, Patryk / Wieczorek, Jan (2024): “Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles”, in: Calzolari, Nicoletta / Kan, Min-Yen / Hoste, Veronique / Lenci, Alessandro / Sakti, Sakriani / Xue, Nianwen (eds.): Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.1500/.
  5. McGovern, Hope / Sirin, Hale / Lippincott, Tom / Caines, Andrew (2024): “Detecting Narrative Patterns in Biblical Hebrew and Greek”, in: Pavlopoulos, John / Sommerschield, Thea / Assael, Yannis et al. (eds.): Proceedings of the 1st Workshop on Machine Learning for Ancient Languages (ML4AL 2024). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.ml4al-1.26.
  6. Ni, Jingwei / Jin, Zhijing / Freitag, Markus / Sachan, Mrinmaya / Schölkopf, Bernhard (2022): “Original or Translated? A Causal Analysis of the Impact of Translationese on Machine Translation Performance”. arXiv preprint. https://doi.org/10.48550/arXiv.2205.02293.
  7. Porter, J. D. / Ilchuk, Yulia / Dombrowski, Quinn (2022): “Linguistic Fingerprints on Translation’s Lens”, in: Journal of Data Mining & Digital Humanities, special issue. https://doi.org/10.46298/jdmdh.7223.
  8. Schleiermacher, Friedrich (2012): “On the Different Methods of Translating”, trans. Susan Bernofsky, in: Venuti, Lawrence (ed.): The Translation Studies Reader. 3rd ed. New York: Routledge, 43–63. First delivered in 1813.
  9. Sirin, Hale / Li, Sabrina / Lippincott, Tom (2024): “Detecting Structured Language Alternations in Historical Documents by Combining Language Identification with Fourier Analysis”, in: Bizzoni, Yuri / Degaetano-Ortlieb, Stefania / Kazantseva, Anna / Szpakowicz, Stan (eds.): Proceedings of the 8th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2024). Association for Computational Linguistics. https://aclanthology.org/2024.latechclfl-1.6/.
  10. Underwood, Ted / Kimutis, Patrick / Witte, Jessica (2020): “NovelTM Datasets for English-Language Fiction, 1700–2009”, in: Journal of Cultural Analytics 5, 2. https://doi.org/10.22148/001c.13147.
  11. Walsh, John A. / Capitanu, Boris / Dubnicek, Ryan / Jett, Jacob / Kudeki, Deren / Layne-Worthey, Glen / Liyanage, Samitha / Organisciak, Peter / Satheesan, Sandeep Puthanveetil / Sepúlveda Torres, Lianet / Swatscheno, Janet / Downie, J. Stephen (2025): The HathiTrust Research Center Extracted Features Dataset. Version 2.5. HathiTrust Research Center. https://doi.org/10.13012/PXP0-F135.