DH 2026

Daejeon, July 27–31

Thu, July 3011:00–12:30S063106
Short Paper

Engagement at Scale: Integrating LLM-Assisted Annotation into TEI Workflows for Digital Cultural Heritage Projects

Joerg Wettlaufer
Academy of Sciences and Humanities in Lower Saxony at Göttingen, Germany · jwettla@gwdg.de

Over the course of nearly fifteen years (2010-2025), the research project Johann Friedrich Blumenbach – Online, conducted at the Academy of Sciences and Humanities in Lower Saxony at Göttingen, has produced one of the most extensive digital editions to date of the works of an Enlightenment-era natural scientist and of his natural history collection. The project set out to digitally interlink Johann Friedrich Blumenbach’s (1752–1840) publications, correspondence, and collection objects, as well as their contemporary and later reception, and to make them accessible through a publicly available online portal.

https://blumenbach-online.de. The TEI P5–encoded texts are available at https://dea-test.gbv.de/. See also the status report of the VZG for 2022, available at https://www.gbv.de/informationen/Verbundzentrale/Publikationen/publikationen-der-vzg-2023/pdf/neumann_231108_mcrworkshop_digitaleeditionen.pdf [12.12.2025].

Beyond its philological and historical ambitions, the project represents a long-term engagement with digital cultural heritage as an evolving methodological practice, in which editorial traditions, data modelling decisions, and technical infrastructures mutually shape one another.

Despite its scope and duration, the project also reveals the structural limits of conventional digital-humanities workflows. Owing to the sheer volume of texts requiring annotation and to the deliberately chosen, highly controlled editorial workflow, not all materials could be semantically encoded to the same depth within the available timeframe.

Of the works listed in the bibliography, 904 out of 1,025 recorded titles have been annotated according to the information provided by the portal.

While the natural history collection objects were recorded almost completely in a structured database environment, other textual components—most notably the numerous editions and translations of Blumenbach’s Handbuch der Naturgeschichte—were encoded only selectively, with full TEI annotation exemplified by a single volume.

Of the 31 editions and translations of the Handbuch der Naturgeschichte, only one copy—no. 27 in the catalogue by C. Kroke—is available in fully encoded TEI P5 form. It is therefore estimated that approximately 20,000 printed pages remain semantically not encoded.

This imbalance highlights a central tension in digital cultural heritage projects: the desire for precision, transparency, and scholarly control on the one hand, and the need for scalability and broader engagement with large and heterogeneous corpora on the other.

This paper critically examines the manual TEI-based annotation workflow established in Blumenbach-Online and situates it in relation to more recent developments in machine learning and Large Language Models (LLMs) for the semantic annotation of historical sources (Christofaro / Zilio 2025, Pollin et al. 2024, Strutz 2026, Galka / Vogeler 2025). The discussion is framed by the concept of engagement at scale: engagement with texts, objects, and historical actors through meticulous scholarly encoding, but also engagement with computational methods that promise to expand the reach and accessibility of digital cultural heritage.

Within the project, annotation was carried out through a predominantly manual workflow based on TEI P5 XML. Structural and semantic encoding relied on expert-driven identification of persons, places, objects, taxonomic terms, and conceptual entities, supported by project-specific guidelines and controlled vocabularies. (Lauer / Weber 2019). This approach ensured a high degree of semantic precision and interpretative reliability, particularly important when dealing with historically variable naming practices, implicit references, and the complex epistemic frameworks of eighteenth-century natural history. At the same time, such deep-level encoding (TEI Level P5) proved extremely labor- and time-intensive. The need for sustained scholarly attention limited the overall recall of annotated entities and relationships across the corpus, effectively constraining the scale of engagement that could be achieved through manual means alone.

In contrast, contemporary machine-learning-based approaches—especially those leveraging LLMs within NLP pipelines—offer new possibilities for large-scale engagement with historical texts. Techniques such as Named Entity Recognition and semantic classification can process extensive and heterogeneous corpora automatically and achieve comparatively high recall in identifying potentially relevant entities and links. For repetitive or formally regular tasks, these methods can significantly reduce editorial workload and accelerate the creation of semantic connections between textual passages and structured object data. An earlier exploratory project, Semantic Blumenbach (2012–2015), already demonstrated the potential of semi-automated approaches by using a curated gazetteer to semantically annotate all German editions of the Handbuch der Naturgeschichte at TEI P5 level.

Cf. http://dhfv-ent2.gcdh.de/blumenbach/ [12.12.2025].

While this method achieved very high precision due to its reliance on a controlled list of entities, it also illustrated the trade-off between coverage and flexibility inherent in rule-based or list-driven approaches (Wettlaufer et al. 2015).

LLM-based methods extend this spectrum by introducing generative and context-sensitive capabilities. However, their strengths in recall and pattern recognition are counterbalanced by risks of overgeneration, false positives, and historically implausible associations. Without careful calibration and scholarly oversight, LLMs may obscure rather than clarify the epistemic structures embedded in historical sources. Precision, transparency, and reproducibility therefore remain critical concerns when integrating such models into editorial workflows (for an evaluation framework see Strutz 2026).

Against this backdrop, the paper discusses several concrete scenarios in which LLMs can be meaningfully integrated into TEI-based workflows as instruments of engaged assistance rather than autonomous annotation. First, LLMs can be employed for pre-annotation of entities and relations, generating suggestions for personal names, place references, and object mentions that are subsequently reviewed and validated by editors. Second, they can support automated classification and ontology mapping, for example in the taxonomic description of natural history objects, provided that the models are constrained by project-specific ontologies and domain knowledge. Third, LLMs can assist in linking textual and object datasets by proposing hypotheses about references across publications, correspondence, and collection records, thereby expanding the space of scholarly inquiry while preserving editorial authority. Finally, it is also possible to attempt to identify places and persons with the help of an LLM and to link them to authority databases such as the GND or the Getty TGN. However, the identification of lesser-known historical figures in particular often remains a challenge and, like the other processing steps, must subsequently be verified carefully. An experiment using OpenAI ChatGPT 5.2 Thinking for the semantic annotation of a German edition of the Handbuch der Naturgeschichte is available at the following link.

See https://chatgpt.com/share/69f30c2a-815c-8384-9b6b-bb1f767aaf84 (in German), [07.05.2026].

Using Blumenbach-Online as a case study, the paper argues that sustainable engagement with digital cultural heritage requires hybrid workflows that combine expert-driven TEI encoding with LLM-based assistance. Such workflows allow projects to scale their engagement with sources without sacrificing the methodological rigor and transparency that underpin scholarly digital editions. More broadly, the contribution positions LLMs as catalysts for rethinking editorial engagement: not as replacements for human interpretation, but as tools that reshape how scholars interact with large corpora, negotiate precision and recall, and mediate cultural heritage to diverse audiences. In doing so, the paper contributes to ongoing debates on “Doing Cultural Heritage” in the digital age by offering a critically grounded, practice-oriented model for engagement at scale.

References
  1. Cristofaro, Marco De / Zilio, Daniel (2025): Automating XML-TEI Encoding of Unpublished Correspondence: A Comparative Analysis of two LLM Approaches, Umanistica Digitale, 531-537. https://iris.unistrasi.it/bitstream/20.500.14091/16941/1/AIUCD2025_De Cristofaro.pdf [07.05.2026]
  2. Galka, Selina / Vogeler, Georg (2025): “Annotating, Projecting, and Interpreting Named Entities in Digital Scholarly Editions with LLMs”, in: International Journal of Digital Humanities 7, 483–509 https://doi.org/10.1007/s42803-025-00114-8
  3. Lauer, Gerhard / Weber, Heiko (2019): “Johann Friedrich Blumenbach – Online.” Johann Friedrich Blumenbach: Race and Natural History, 1750–1850, edited by Nikolaas Rupke and Gerhard Lauer, Routledge, 16–24. (Routledge Studies in the History of Science, Technology and Medicine, 1).
  4. Pollin, Christopher / Fischer, Franz / Sahle, Patrick / Scholger, Martina / Vogeler, Georg (2025): “When It Was 2024 – Generative AI in the Field of Digital Scholarly Editions” in: Zeitschrift für digitale Geisteswissenschaften, 10. https://doi.org/10.17175/2025_008
  5. Strutz, Sabrina (2026): “A Multi Dimensional Evaluation Framework for Assessing LLM Performance in TEI Encoding”, in: Journal of Open Humanities Data, 12: 39, 1–20. DOI: https://doi.org/10.5334/johd.484
  6. Wettlaufer, Jörg / Johnson, Christopher H. / Scholz, Martin / Fichtner, Mark / Thotempudi, Sree G. (2015): “Semantic Blumenbach: Exploration of Text–Object Relationships with Semantic Web Technology in the History of Science.” Digital Scholarship in the Humanities, 30, suppl. 1, i187–i198. Special Issue: Digital Humanities 2014, edited by Melissa Terras et al., https://academic.oup.com/dsh/article/30/suppl_1/i187/364720/[07.05.2026].