Daejeon, July 27–31
Over the course of nearly fifteen years (2010-2025), the research project Johann Friedrich Blumenbach – Online, conducted at the Academy of Sciences and Humanities in Lower Saxony at Göttingen, has produced one of the most extensive digital editions to date of the works of an Enlightenment-era natural scientist and of his natural history collection. The project set out to digitally interlink Johann Friedrich Blumenbach’s (1752–1840) publications, correspondence, and collection objects, as well as their contemporary and later reception, and to make them accessible through a publicly available online portal. https://blumenbach-online.de. The TEI P5–encoded texts are available at https://dea-test.gbv.de/. See also the status report of the VZG for 2022, available at https://www.gbv.de/informationen/Verbundzentrale/Publikationen/publikationen-der-vzg-2023/pdf/neumann_231108_mcrworkshop_digitaleeditionen.pdf [12.12.2025].
Despite its scope and duration, the project also reveals the structural limits of conventional digital-humanities workflows. Owing to the sheer volume of texts requiring annotation and to the deliberately chosen, highly controlled editorial workflow, not all materials could be semantically encoded to the same depth within the available timeframe. Of the works listed in the bibliography, 904 out of 1,025 recorded titles have been annotated according to the information provided by the portal. Of the 31 editions and translations of the Handbuch der Naturgeschichte, only one copy—no. 27 in the catalogue by C. Kroke—is available in fully encoded TEI P5 form. It is therefore estimated that approximately 20,000 printed pages remain semantically not encoded.
This paper critically examines the manual TEI-based annotation workflow established in Blumenbach-Online and situates it in relation to more recent developments in machine learning and Large Language Models (LLMs) for the semantic annotation of historical sources (Christofaro / Zilio 2025, Pollin et al. 2024, Strutz 2026, Galka / Vogeler 2025). The discussion is framed by the concept of engagement at scale: engagement with texts, objects, and historical actors through meticulous scholarly encoding, but also engagement with computational methods that promise to expand the reach and accessibility of digital cultural heritage.
Within the project, annotation was carried out through a predominantly manual workflow based on TEI P5 XML. Structural and semantic encoding relied on expert-driven identification of persons, places, objects, taxonomic terms, and conceptual entities, supported by project-specific guidelines and controlled vocabularies. (Lauer / Weber 2019). This approach ensured a high degree of semantic precision and interpretative reliability, particularly important when dealing with historically variable naming practices, implicit references, and the complex epistemic frameworks of eighteenth-century natural history. At the same time, such deep-level encoding (TEI Level P5) proved extremely labor- and time-intensive. The need for sustained scholarly attention limited the overall recall of annotated entities and relationships across the corpus, effectively constraining the scale of engagement that could be achieved through manual means alone.
In contrast, contemporary machine-learning-based approaches—especially those leveraging LLMs within NLP pipelines—offer new possibilities for large-scale engagement with historical texts. Techniques such as Named Entity Recognition and semantic classification can process extensive and heterogeneous corpora automatically and achieve comparatively high recall in identifying potentially relevant entities and links. For repetitive or formally regular tasks, these methods can significantly reduce editorial workload and accelerate the creation of semantic connections between textual passages and structured object data. An earlier exploratory project, Semantic Blumenbach (2012–2015), already demonstrated the potential of semi-automated approaches by using a curated gazetteer to semantically annotate all German editions of the Handbuch der Naturgeschichte at TEI P5 level. Cf. http://dhfv-ent2.gcdh.de/blumenbach/ [12.12.2025].
LLM-based methods extend this spectrum by introducing generative and context-sensitive capabilities. However, their strengths in recall and pattern recognition are counterbalanced by risks of overgeneration, false positives, and historically implausible associations. Without careful calibration and scholarly oversight, LLMs may obscure rather than clarify the epistemic structures embedded in historical sources. Precision, transparency, and reproducibility therefore remain critical concerns when integrating such models into editorial workflows (for an evaluation framework see Strutz 2026).
Against this backdrop, the paper discusses several concrete scenarios in which LLMs can be meaningfully integrated into TEI-based workflows as instruments of engaged assistance rather than autonomous annotation. First, LLMs can be employed for pre-annotation of entities and relations, generating suggestions for personal names, place references, and object mentions that are subsequently reviewed and validated by editors. Second, they can support automated classification and ontology mapping, for example in the taxonomic description of natural history objects, provided that the models are constrained by project-specific ontologies and domain knowledge. Third, LLMs can assist in linking textual and object datasets by proposing hypotheses about references across publications, correspondence, and collection records, thereby expanding the space of scholarly inquiry while preserving editorial authority. Finally, it is also possible to attempt to identify places and persons with the help of an LLM and to link them to authority databases such as the GND or the Getty TGN. However, the identification of lesser-known historical figures in particular often remains a challenge and, like the other processing steps, must subsequently be verified carefully. An experiment using OpenAI ChatGPT 5.2 Thinking for the semantic annotation of a German edition of the Handbuch der Naturgeschichte is available at the following link. See https://chatgpt.com/share/69f30c2a-815c-8384-9b6b-bb1f767aaf84 (in German), [07.05.2026].
Using Blumenbach-Online as a case study, the paper argues that sustainable engagement with digital cultural heritage requires hybrid workflows that combine expert-driven TEI encoding with LLM-based assistance. Such workflows allow projects to scale their engagement with sources without sacrificing the methodological rigor and transparency that underpin scholarly digital editions. More broadly, the contribution positions LLMs as catalysts for rethinking editorial engagement: not as replacements for human interpretation, but as tools that reshape how scholars interact with large corpora, negotiate precision and recall, and mediate cultural heritage to diverse audiences. In doing so, the paper contributes to ongoing debates on “Doing Cultural Heritage” in the digital age by offering a critically grounded, practice-oriented model for engagement at scale.