DH 2026

Daejeon, July 27–31

Fri, July 3114:00–15:30S106204-205
Short Paper

Keeping Uncertainty Visible: Engagement and Ethics in AI-Supported Provenance Research

Elisa Ludwig
Museum Ulm, Germany · e.ludwig@ulm.de
Toni Thomä
1tmsolutions · t.thomae@1tm.solutions
Martin Schubert
1tmsolutions · m.schubert@1tm.solutions

1. Provenance Sources: Infrastructure and Limitations

Provenance research relies on archival sources to reconstruct the ownership histories of cultural objects. The growing importance of due diligence and traceable sources is driving large-scale digital infrastructures that have further advanced provenance research (e.g., Getty Provenance Index, Heidelberg Accession Index, and German Sales). GLAM institutions struggle with the development and long-term maintenance of their own digital research data infrastructures, which require technical expertise, financial investment, and sustained institutional support conditions. (Hopp 2018, Hopp/von dem Bussche 2023, Sepp 2023; Schneider et al. 2026). The Ulm project addresses these gaps and seeks a source-critical, FAIR-compliant database system within realistic resource constraints.

2. The Director’s Correspondence (1933-1945): Small Data as a Productive Epistemic Condition

The nearly completely preserved collection of Director’s Records of the Museum Ulm (1933-1945) comprises 25,000 documents, enabling longitudinal analysis of institutional decision-making during the National Socialist period in Germany. The records document interactions among international art markets, cultural institutions and private
collector.collector.



Fig. 1: Examples from the „Director’s Records of the Museum Ulm: right: A letter by Oskar Schlemmer to Julius Baum (13.02.1933, MUDA-03-B- 285 and left: Letter to Adolf Häberle to Museum Ulm (1937), MUDA-03-E-102b.

As a result, the source material is semantically heterogeneous, featuring varied document types, correspondence, handwriting systems and languages. Such datasets are particularly important for testing AI systems in less-studied languages and specialised application areas, and make it easier to analyse whether the AI is biased, how it makes decisions, and how errors propagate within the system. By treating this data as a historical and social artefact, we want to highlight the selection, annotation, and modelling process, making scholarly interpretation and its uncertainty transparent.As a result, the source material is semantically heterogeneous, featuring varied document types, correspondence, handwriting systems and languages. Such datasets are particularly important for testing AI systems in less-studied languages and specialised application areas, and make it easier to analyse whether the AI is biased, how it makes decisions, and how errors propagate within the system. By treating this data as a historical and social artefact, we want to highlight the selection, annotation, and modelling process, making scholarly interpretation and its uncertainty transparent.

3. AI Across Transcription, Structuring, and Publication

In addition to an in-depth exploration of current research discussions and solution strategies, we interviewed professionals and stakeholders from related projects in archives (Brandenburg State Archives, Baden-Württemberg State Archives), museums (The B&O Railroad Museum, Baden-Württemberg State Museum), and libraries (Heidelberg University Library), and consulted experts from data competence centres (NFDI4Culture, text+, SODa) to identify solutions for the different processes.


To address the challenges named above, the project develops a five-stage modular AI-supported data pipeline with various tools designed for small, resource- constrained institutions. Methodologically, we used prompt design, annotation strategies, and model evaluation processes to assess and improve data quality. The enriched data is mapped to the Linked. Art Model, which is based on the CIDOC-CRM ontology. The architecture deliberately separates three data layers and is set up as a modular graph that separates documentary evidence, interpretive assertions, and authority control. A raw layer holds a machine-readable re-presentation of the document, typically an image embedding and an optional structured JSON, generic and not yet bound to any institutional schema. A structured layer carries the institution-specific extraction, along with field-level confidence scores, and a layer-specific embedding that powers the hybrid search (vector plus full-text). An enriched layer links the structured data to external vocabularies such as Getty AAT/ULAN, GND, and WikiData via API. Human-in-the-loop corrections at the structured layer trigger re-computation of the corresponding embedding so that the hybrid search reflects the corrected state; older versions remain stored and com-parable but are not surfaced in retrieval. This layered, versioned design is what makes the historical-critical method operationally implementable and what keeps every step inspectable, revisable, and documentable.To address the challenges named above, the project develops a five-stage modular AI-supported data pipeline with various tools designed for small, resource- constrained institutions. Methodologically, we used prompt design, annotation strategies, and model evaluation processes to assess and improve data quality. The enriched data is mapped to the Linked. Art Model, which is based on the CIDOC-CRM ontology. The architecture deliberately separates three data layers and is set up as a modular graph that separates documentary evidence, interpretive assertions, and authority control. A raw layer holds a machine-readable re-presentation of the document, typically an image embedding and an optional structured JSON, generic and not yet bound to any institutional schema. A structured layer carries the institution-specific extraction, along with field-level confidence scores, and a layer-specific embedding that powers the hybrid search (vector plus full-text). An enriched layer links the structured data to external vocabularies such as Getty AAT/ULAN, GND, and WikiData via API. Human-in-the-loop corrections at the structured layer trigger re-computation of the corresponding embedding so that the hybrid search reflects the corrected state; older versions remain stored and com-parable but are not surfaced in retrieval. This layered, versioned design is what makes the historical-critical method operationally implementable and what keeps every step inspectable, revisable, and documentable.

Each AI processing run is versioned and stored alongside its prompt template and model identifier, making the inference chain transparent and reproducible across iterations. The same architectural mechanism that supports correction also enables systematic comparison of alternative VLM and LLM configurations against a gold standard sample, without requiring redesign of the data model. Hereby we are taking into account the challenges identified by Holle Meding and Aurel Daug of LLMs in historical research (Meding & Daugs 2026). Confidence scores from all steps are propagated, allowing uncertainty to be modelled throughout. The central research question is how AI-supported workflows can explicitly model and preserve uncertainty, vagueness, and subjectivity in provenance data to promote ethical engagement and transparency.

4. Modelling Uncertainties of Archival Provenance Data

Building on research by Smets (1997), Nagypál/Motik, Piotrowski (2019), and Mariani (2022), the goal is to combine solutions and account for vagueness, incompleteness, subjectivity, and their impact on uncertainty in archival sources (Andratschke et al. 2019, Lang 2023, Ginhoven 2022, Barget 2025 and Battisti 2025). Drawing on Mariani’s four epistemic frontiers, we extend VISU-based approaches through critical source analysis. The discussions with the ProvEnhance team, based at the Museum of Fine Arts in Brussels, provided valuable insights for the data collection tables and for modelling uncertainties (Leroux/Almstadt 2026). In line with König’s concept of epistemic incompleteness, our method aims to transparently document uncertainties, gaps, and parallel argumentations, treating them as provisional yet essential for scholarly engagement (König 2026).

We address algorithmic bias, historical sensitivity, and legal compliance by evaluating potential amplification of bias. Using technical control mechanisms (such as robots.txt, the X-Robots-Tag, and the TDM Reservation Protocol), we wanted to frame discriminatory language and historical contexts visible (Zimmermann/Dau 2025). Data protection and redaction align with European regulations (EU AI Act). Human-in-the-loop validation provides oversight at all stages and ensures dataset quality.

5. Conclusion

Showcasing the tension between theory and practice, this paper reflects on whether combining small-data epistemology, AI-supported workflows, and interoperable semantic modelling can create a scalable, transparent approach to provenance research under real-world constraints, which supports interoperability within GLAM institutions and external infrastructures. In the context of DH2026’s theme Engagement, AI is not positioned as a black-box solution but as a transparent, critically examined partner in the production of historically situated knowledge. By sharing the considerations behind tool selection and development processes with other institutions, we hope to obtain constructive feedback to further develop the data pipeline and want to offer ideas for other museums to engage with their sources, their data, and their communities.

References
  1. Andratschke, C. et al. (2019): Leitfaden zur Standardisierung von Provenienzangaben, Berlin. <http://arbeitskreisprovenienzfor-schung.org/data/uploads/Leitfaden_PFeV_online.pdf> [05.05.2026].
  2. Barget, M. (2024): „Nicht-Wissen modellieren & visualisieren - interdisziplinäre Perspektiven auf Datenlücken & unsichere Daten“, in: Ringvorlesung Digitale Methoden, Mainz. DOI: https://doi.org/10.5281/zenodo.11237601.
  3. Battisti, T. / Daquino, M. (2025): “Uncertainty, narrativity, and critical approaches in Digital Humanities information visualisation projects”, in: Proceedings of the 21st Conference on Information and Research science Connecting to Digital and Library science, <https://hdl.handle.net/11585/1009129> [05.05.2026].
  4. Hopp, M. (2018): “Provenienzrecherche und digitale Forschungsinfrastrukturen in Deutschland: Tendenzen, Desiderate, Bedürfnisse“, in: Blimlinger, Eva / Schödl, Heinz (eds.): …(k)ein Ende in Sicht. 20 Jahre Kunstrückgabegesetz in Österreich. Wien: 35–59.
  5. Hopp, M. / van dem Bussche, R. (2023): „Provenienzforschung und ihre Quellenbestände. Aktuelle Nutzungsszenarien zwischen Open Access und Inaccessibility“, in: Helling, P., Majka, N., & Gius, E. (ed.): Open Humanities. Open Cultures (DHd Conference 2023), Trier/Belval, March 2023: 14. DOI: https://doi.org/10.5281/zenodo.10977717.
  6. König, M. (2026): “Fertig – vorerst. Unfertigkeit als epistemischer Wert in den digitalen Geisteswissenschaften“, in: Zeitschrift für digitale Geisteswissenschaften 11 (05.02.2026). DOI: 10.17175/2026002.
  7. Kresmer, A. / Goldman, J. (2025): “Unlocking the Past with AI at the B&O Railroad Museum”, in: American Alliance Museum Magazine June/July 2025: 32-37. <https://www.aam-us.org/2025/06/29/unlocking-the-past-with-ai-at-the-bo-railroad-museum/> [05.05.2026].
  8. Lang, S. (2023): „Immer FAIR?! Problematische Inhalte in den Datenbeständen der Provenienzforschung“, in: Helling, Patrick (ed.): FORGE 2023 - Anything Goes?! Forschungsdaten in den Geisteswissenschaften - kritisch betrachtet, Book of Abstracts, Tübingen, October 2023: 84–92. DOI: https://doi.org/10.5281/zenodo.15227793.
  9. Lincoln, M., van Ginhoven, S. (2020): “Modeling a Fragmented Archive: A Missing Data Case Study from Provenance Research”, in: ADHO (eds.): Puentes/Bridges ADHO June 2018: DOI: https://doi.org/10.1184/R1/12363059.v1.
  10. Leroux, A., & Almstadt, F. (2026): “Supplementary Material - From Objects to the Art Market. A data model for provenance research at the Royal Museums of Fine Arts of Belgium, 1933–1960 (1.0)“ [Data set]. DOI: https://doi.org/10.5281/zenodo.19822067.
  11. Mariani, F. (2022): “Introducing VISU: Vagueness, Incompleteness, Subjectivity, and Uncertainty in art provenance data”, in: Rochat, Yannik et al. (eds): COMHUM 2022 Computational Methods in the Humanities 2022: proceedings of the Workshop on Computational Methods in the Humanities 2022, Lausanne, June 2022: 63–84. <https://ceur-ws.org/Vol-3602/paper5.pdf> [05.05.2026].
  12. Nagypál, G. / Motik, B. (2003): “A Fuzzy Model for Representing Uncertain, Subjective, and Vague Temporal Knowledge in Ontologies“, in: Meersman, R. / Tari, Z., Schmidt, D.C. (eds.): On The Move to Meaningful Internet Systems 2003: CoopIS, DOA, and ODBASE. OTM 2003. Lecture Notes in Computer Science, vol. 2888. Berlin, Heidelberg. DOI: https://doi.org/10.1007/978-3-540-39964-3_57.
  13. Meding, H. and Daugs, A. (2026): “LLMs in den Geschichtswissenschaften: Potenziale, Grenzen und Anwendungs­beispiele (NER & RAG)”, in: DHd-AG Angewandte Generative KI in den Digitalen Geisteswissenschaften (13.04.2026) <https://lisa.gerda-henkel-stiftung.de/llms_in_den_geschichtswissenschaften> [05.05.2026].
  14. Piotrowski, M. (2019): “Accepting and Modeling Uncertainty“, in: Kuczera, Andreas et al. (eds.): Die Modellierung des Zweifels – Schlüsselideen und -konzepte zur graphbasierten Modellierung von Unsicherheiten. Wolfenbüttel: DOI: 10.17175/sb004_006a.
  15. Rother, L. et al. (2022): “Taking Care of History: Toward a Politics of Provenance Linked Open Data in Museums”, in: L. Fry und E. Canning (eds.): Perspectives on Data. Chicago <https://www.artic.edu/digital-publications/37/perspectives-on-data/25/taking-care-of-history-toward-a-politics-of-provenance-linked-open-data-in-museums> [05.05.2026].
  16. Smets, P. (1997): „Imperfect Information: Imprecision and Uncertainty“, in: Motro, A. / Smets, P. (eds.): Uncertainty Management in Information Systems. Boston: 225–254. DOI: https://doi.org/10.1007/978-1-4615-6245-08.
  17. Schneider, S. et. al. (2026): “From origin to future: interdisciplinary approaches to provenance in art collections“, in: Museum Management and Curatorship: 1–24. https://doi.org/10.1080/09647775.2026.2654871.
  18. Sepp, T. (2023): “Bedingt gute Rahmenbedingungen. Zur digitalen Provenienzforschung“, in: Kunstchronik 76, 7: 363–370.