DH 2026

Daejeon, July 27–31

Thu, July 3016:30–18:00S107204-205
Short Paper

TextPloring: Expanding (Meta)data Modeling for Cross-Disciplinary Data Exploration and Reuse

Leo Born
Humboldt-Universität zu Berlin, Germany · leo.born@hu-berlin.de
Paula Rebecca Schreiber
Humboldt-Universität zu Berlin, Germany · paula.rebecca.schreiber@hu-berlin.de
Isabell Trilling
Humboldt-Universität zu Berlin, Germany · isabell.trilling.1@hu-berlin.de
Maik Bierwirth
Humboldt-Universität zu Berlin, Germany · maik.bierwirth@hu-berlin.de
Toska Brandl
Humboldt-Universität zu Berlin, Germany · toska.brandl.1@hu-berlin.de
Malte Dreyer
Humboldt-Universität zu Berlin, Germany · dreyer@hu-berlin.de
Torsten Hiltmann
Humboldt-Universität zu Berlin, Germany · torsten.hiltmann@hu-berlin.de
Anke Lüdeling
Humboldt-Universität zu Berlin, Germany · anke.luedeling@hu-berlin.de
Tsimafei Luzan
Humboldt-Universität zu Berlin, Germany · tsimafei.luzan.1@hu-berlin.de
Carolin Odebrecht
Humboldt-Universität zu Berlin, Germany · carolin.odebrecht@hu-berlin.de

Extending a Research Data Repository

TextPloring is a project demonstrating the adaptability of metadata-driven access to research data and bridging the gap in inter-disciplinary data reuse. To that end, this short paper reports on our work to extend the scope of LAUDATIO (Guescini / Odebrecht 2021) to the field of history. LAUDATIO is a research data repository for deeply annotated corpora established in the context of historical linguistics and currently hosts dozens of corpora that, while annotated from a linguistic viewpoint, would provide valuable research data for historians as well. The underlying metadata model (Odebrecht 2018) – based on an ODD specification of TEI with a subset of TEI elements, modules, and attributes – is sufficiently generic to support broader application across historically oriented disciplines. It is realised as TEI-XML, using the teiHeader to define a corpus and its constituents (documents and annotations) and to capture both internal information (e.g. what it contains) as well as external information (e.g. who compiled it and which technologies were used). Specifically, the project aims to facilitate the enrichment of linguistically annotated corpora with historically informed annotations to reshape the opportunities afforded by the composition of the data. The publication of updated corpus versions will enable these enhancements to open up new avenues of research across disciplines. Due to a shared need for textual data, increasing the number of corpora accessible beyond their (respective) disciplines and providing greater explorability through LAUDATIO ultimately benefits linguists and historians alike.

While existing textual research data repositories have made substantial contributions by focusing on specific time periods (e.g. Deutsches Textarchiv (Berlin-Brandenburgische Akademie der Wissenschaften 2026) covers data from the 17th to the 20th century), modalities (e.g. HZSK (Jettka / Stein 2014), which primarily comprises spoken language), or genres (e.g. correspSearch (Dumont 2016), a web resource for scholarly editions of letters), we envision LAUDATIO as a complementary repository designed to accommodate a diverse array of language data across time, modality, and form. Given that LAUDATIO already adheres to FAIR data principles (Wilkinson et al. 2016), we place substantial emphasis on ensuring the findability and reusability of its data. Accordingly, this short paper reports on expert interviews conducted with scholars from history and linguistics to ascertain key areas in which the current metadata model requires enhancement or adaptation. Expert interviews constitute an effective approach to gain a high level of contextual knowledge essential for domain-specific metadata modeling (Reinhold 2015). For each discipline, we selected ten participants using a purposive sampling strategy designed to capture the methodological and epistemic diversity of the field. Selection criteria included subfield coverage, methodological profile, use of metadata in the research workflow, and engagement with digital and data-driven approaches, balanced across career stages: early-career (< 5 years), mid-career (5-10 years), and senior researchers (> 10 years). The interviews were preceded by a short questionnaire as part of a two-stage design that combines quantitative pre-assessment with an approximately 30-minute qualitative interview.

In conjunction with these expert interviews, we conduct user studies to identify user needs and emerging research possibilities through a bottom-up approach. These studies focus on existing and newly prototyped functionalities for data exploration and searchability, given their central role in facilitating engagement with complex research datasets. The studies are set up at multiple stages of the project: internal workshops with direct project members (e.g. students, associated researchers) are held at regular intervals, enabling an agile feedback and development loop. Complementary to these are external workshops (that we either host independently or as satellite workshops of conferences) where researchers are tasked to use LAUDATIO and determine its feasibility in aiding data exploration. The studies are based on heuristic walkthroughs (Sears 1997), a two-step process wherein we first provide guided tasks that need to be completed, followed by a free assessment of the user interface. These walkthroughs are informed by a custom set of usability heuristics (Stiller et al. 2016, Rodwell et al. 2025), such as adherence to conventions and expectations regarding naming and user interactions, sensible information architectures, and disciplinary inclusivity. Results are collected in-person or online by completing task sheets and optionally recording the screen of the test sessions.

Visual-Driven Exploration of Research (Meta)data

Visualization of research (meta)data in this context serves to support data interpretation and to improve the accessibility of the resulting knowledge. We thus aim to extend LAUDATIO with effectively designed (Franconeri et al. 2021), modular visualizations that advance data exploration and enable researchers to query the data according to their own research questions (cf. Figures 1-3 that show three prototyped modes of visualizing metadata for an exploratory access to LAUDATIO across corpora). The selection of research data can be guided by a multitude of factors.

Figure Displaying corpus-level characteristics, here document date ranges.
Figure Aggregate visualizations, here the distribution of languages.
Figure Network views that connect corpora if they share a language, for example.

One relevant question when exploring a corpus concerns how representative its constituent documents are. Formal representation can, for example, be measured by the extent to which data has been collected along the dimensions of breadth (i.e. the number of documents per year) and depth (i.e. the length of documents per year). Visualizations can readily answer such a question (cf. Figures 4, 5), thus determining the suitability of a corpus for certain research avenues (e.g. those where depth is considered more important than breadth) while simultaneously enabling insight into the composition of the data. These considerations are based on concrete requirements drawn from user stories developed in conjunction with the relevant research communities (Althage et al. 2022).

Figure The distribution of documents over time (the number of documents per year).
Figure The distribution of tokens over time (the aggregate number of tokens per year). Note the peak in 1870 to the right, indicating that while there are only few documents (two, cf. Figure 4) for that year, they are much longer than documents of years with similarly few documents.

Informed by these user stories and as a preparatory step for the user studies, we also improved the repository’s usability across disciplinary contexts and implemented new functions for the automatic extraction, display, and filtering of plaintext data, enhancing the dimensionality of the accessible data (cf. Figures 6, 7).

Figure The extraction of plaintext data allows for more downstream exploratory possibilities. The current development version allows the additional contents to be viewed, downloaded, and searched. Shown is one document of the Mastro – Mittelalterliche Stadtchroniken corpus.

Figure The new plaintext search interface displays results and highlights example matches in context. Furthermore, users can download all found documents at once. Shown is the search results page for the query "Beifuß" (German for "mugwort").

The user studies draw on the actual needs of (re)users of research data and inform the implementation of easy-to-use exploration mechanisms for the corpora in LAUDATIO. Particular emphasis is placed on two tangible data use cases:

  • A new version of the RIDGES corpus, a diachronic corpus of German herbal texts from the 15th to the 20th century (Odebrecht et al. 2017, Uyanık et al. 2025)
  • The addition of newly digitized texts from the Mastro – Mittelalterliche Stadtchroniken volumes, a collection of medieval city chronicles for 18 different German-speaking cities (Hiltmann / Odebrecht 2023)

These two data use cases are of particular relevance to linguistic and historical scholarship, respectively. They provide the substantive basis for validating the metadata schema extension and for assessing the feasibility of cross-disciplinary data reuse, all while maintaining the integrity of comprehensive disciplinary annotations. We regard this to be a crucial step for the strengthening of inter-disciplinary engagement by crossing institutional borders. We thus strive towards the structural connection of LAUDATIO to the relevant national research data consortia such as NFDI4Memory and Text+. This will be fertile ground for opening up the datasets to a larger user base within the historical and linguistic research communities and further consolidate LAUDATIO as a cross-disciplinary research data repository. Coupled with the visual approach to data exploration, this will also contribute to greater engagement with research data itself, empowering researchers to find and use data suited for their research needs.

Acknowledgements

We would like to thank the anonymous reviewers for their helpful comments. This work has been funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under project number 536177165.

References
  1. Althage, Melanie / Dreyer, Malte / Guescini, Rolf / Hiltmann, Torsten / Lüdeling, Anke / Odebrecht, Carolin (2022): “Referenzierung des digitalen kulturellen (Text-)Erbes – Digitale Quellenkritik und Modellierung von Metadaten”, in: Book of Abstracts – Dhd 2022, March 2022. DOI: 10.5281/zenodo.6322547.
  2. Berlin-Brandenburgische Akademie der Wissenschaften (ed.) (2026): Deutsches Textarchiv. Grundlage für ein Referenzkorpus der neuhochdeutschen Sprache. <https://www.deutschestextarchiv.de/> [04.05.2026].
  3. Dumont, Stefan (2016): “correspSearch – Connecting Scholarly Editions of Letters”, in: Journal of the Text Encoding Initiative 10. DOI: 10.4000/jtei.1742.
  4. Franconeri, Steven L. / Padilla, Lace M. / Shah, Priti / Zacks, Jeffrey M. / Hullman, Jessica (2021): “The Science of Visual Data Communication: What Works”, in: Psychological Science in the Public Interest 22, 3: 110-161. DOI: 10.1177/15291006211051956.
  5. Guescini, Rolf / Odebrecht, Carolin (2021): Laudatio Repository - Long-term Access and Usage of Deeply Annotated Information. Version 1.0rc. DOI: 10.5281/zenodo.5101366.
  6. Hiltmann, Torsten / Odebrecht, Carolin (2023): Transkriptions- und Zoning-Guidelines für das Korpus MaStro (Mittelalterliche Stadtchroniken) Version 1.0. DOI: 10.5281/zenodo.8128664.
  7. Jettka, Daniel / Stein, Daniel (2014): “The HZSK Repository: Implementation, Features, and Use Cases of a Repository for Spoken Language Corpora”, in: D-Lib Magazine 20, 9/10. DOI: 10.1045/september2014-jettka.
  8. Odebrecht, Carolin (2018): MKM – ein Metamodell für Korpusmetadaten. Ph.D. thesis, Humboldt-Universität zu Berlin.
  9. Odebrecht, Carolin / Belz, Malte / Zeldes, Amir / Lüdeling, Anke / Krause, Thomas (2017): “RIDGES Herbology: designing a diachronic multi-layer corpus”, in: Language Resources and Evaluation 51, 3: 695-725. DOI: 10.1007/s10579-016-9374-3.
  10. Reinhold, Anke (2015): “Das Experteninterview als zentrale Methode der Wissensmodellierung in den Digital Humanities”, in: Information - Wissenschaft & Praxis 66, 5-6: 327-333. DOI: 10.1515/iwp-2015-0057.
  11. Rodwell, Elizabeth / Neumann, Kristina / Lindner, Peggy (2025): “User Experience (UX) Heuristics for the Digital Humanities”, in: Digital Humanities Quarterly 19, 4. DOI: 10.63744/ah2t7jpfg4rt.
  12. Sears, Andrew (1997): “Heuristic Walkthroughs: Finding the Problems Without the Noise”, in: International Journal of Human–Computer Interaction 9, 3: 213-234.
  13. Stiller, Juliane / Gnadt, Timo / Romanello, Matteo / Thoden, Klaus (2016): “Anforderungen ermitteln, Lösungen evaluieren und Erfolge messen – Begleitforschung in DARIAH-DE”, in: Bibliothek Forschung und Praxis 40, 2: 250-258. DOI: 10.1515/bfp-2016-0025.
  14. Uyanık, Esra / Müller, Sven Oliver / Lüdeling, Anke / Krause, Thomas (2025): “Differenzierung und Standardisierung. Zur Entwicklung von Registern”, in: Pawłowski, Grzegorz / Guławska, Małgorzata / Bąk, Paweł (eds.): Historische Fach- und Wissenschaftstexte kontrastiv. Berlin: De Gruyter 193-214.
  15. Wilkinson, Mark D. / Dumontier, Michel / Aalbersberg, IJsbrand Jan / Appleton, Gabrielle / Axton, Myles / Baak, Arie / Blomberg, Niklas / Boiten, Jan-Willem / Bonino da Silva Santos, Luiz / Bourne, Philip E. / Bouwman, Jildau / Brookes, Anthony J. / Clark, Tim / Crosas, Mercè / Dillo, Ingrid / Dumon, Olivier / Edmunds, Scott / Evelo, Chris T. / Finkers, Richard / Gonzalez-Beltran, Alejandra / Gray, Alasdair J.G. / Groth, Paul / Goble, Carole / Grethe, Jeffrey S. / Heringa, Jaap / ’t Hoen, Peter A.C / Hooft, Rob / Kuhn, Tobias / Kok, Ruben / Kok, Joost / Lusher, Scott J. / Martone, Maryann E. / Mons, Albert / Packer, Abel L. / Persson, Bengt / Rocca-Serra, Philippe / Roos, Marco / van Schaik, Rene / Sansone, Susanna-Assunta / Schultes, Erik / Sengstag, Thierry / Slater, Ted / Strawn, George / Swertz, Morris A. / Thompson, Mark / van der Lei, Johan / van Mulligen, Erik / Velterop, Jan / Waagmeester, Andra / Wittenburg, Peter / Wolstencroft, Katherine / Zhao, Jun / Mons, Barend (2016): “The FAIR Guiding Principles for scientific data management and stewardship”, in: Scientific Data 3, 160018. DOI: 10.1038/sdata.2016.18.