DH 2026

Daejeon, July 27–31

Wed, July 2909:00–10:30S110108
Short Paper

VLM’s as a Complete Digital Edition Pipeline

Mart Herman Gerrit Makkink
University of Vienna, Austria · herman.gerrit.makkink@univie.ac.at

The field of premodern ethnic identities has long been dominated by nationalist narratives.

Andrew Gillett, “Ethnogenesis: A Contested Model of Early Medieval Europe,” History Compass 4, no. 2 (March 1, 2006): 199.

Often, the insular focus on a specific time or nation prevented a longue durée conceptualisation of representations of ethnic identities in the sources.

Fernand Braudel, La Méditerranée et le monde méditerranéen à l’époque de Philippe II, (Armand Colin, 1966).

While recently such conceptualisations have appeared, their scopes have been limited due to methodological constraints. By working with only a select few sources, scholars of medieval ethnic identity still tend to limit themselves to a conceptualisation of one single ethnic identity (for example Robert Flierman and the Saxons), or one single source (for example Helmut Reimitz and the Fredigar Chronicle).

Robert Flierman, Saxon Identities, AD 150-900 (Bloomsbury Academic, an imprint of Bloomsbury Publishing Plc, 2017). Robert Flierman, Saxon Identities, AD 150-900 (Bloomsbury Academic, an imprint of Bloomsbury Publishing Plc, 2017); Helmut Reimitz, History, Frankish Identity and the Framing of Western Ethnicity, 550-850 (Cambridge University Press, 2015).

Larger meta-analyses are usually based on a reading of secondary literature, rather than re-evaluations of the primary sources themselves.

Patrick J. Geary, The Myth of Nations: The Medieval Origins of Europe (Princeton University Press, 2002); Visions of Community in the Post-Roman World: The West, Byzantium and the Islamic World, 300-1100, ed. Walter Pohl et al. (Ashgate, 2012).

However, to truly step out of a nationalist frame and to escape the biases of a single source on ethnic identification, widening the scope is essential. Rather than focus on one people, which results in a conceptualisation of the ethnic identity of that people, or within one source, which results in a conceptualisation of ethnic identity within that source, a plurality of peoples and sources must be studied.

For this reason, I am looking at the conceptualisation of ethnic identity in European medieval historiography through distant reading methods, which allow for a wider comparison across time and place. The main issue with this approach is that currently, not all source texts exist in a digitised edition. This limits the true potential scope of the project to already published digital editions, with all the biases that come from such a selection.

For example those available on Corpus Corporum: ‘Corpus Corporum’, accessed 15 December 2025, https://mlat.uzh.ch/.

To escape this silo, I am looking at alternative ways to digitise texts and turn them into usable digital editions on a larger scale. Most methods so far have proven time consuming, labour intensive and prone to error, texts created fully with Optical Character Recognition (OCR) need to be cleaned up to get rid of footnotes or to integrate alternative readings. Streamlining this process would greatly increase digitisation of texts, which has wide-ranging benefits beyond this project.

In this paper, I will show a new approach create a pipeline which takes existing scans of print editions and perform OCR, layout detection and data modelling. Using QwenLM, a Vision Language Model (VLM), to create fully marked-up digital editions, encoded in TEI. LLM’s have are known to mostly excel at data conversion tasks.

Xiangru Tang et al., ‘Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?’, arXiv:2309.08963, preprint, arXiv, 4 April 2024, https://doi.org/10.48550/arXiv.2309.08963.

However, for OCR tasks, they are more prone to hallucination.

Zhentao He et al., ‘Seeing Is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models’, arXiv:2506.20168, preprint, arXiv, 22 September 2025, https://doi.org/10.48550/arXiv.2506.20168.

VLM’s take images as input, instead of pure texts, and are easier to finetune for specific image-related tasks. In this paper, I will show that this means that VLM’s could therefore combine the image-related task of OCR with the LLM’s excellence at data conversion to create usable editions of existing texts.

To show the efficacy of this usecase I am using a corpus of texts for which already some TEI editions exist: the dMHG. The entirety of the print Monumenta Historia Germania editions have been scanned and made available online.

dMGH, accessed 15 December 2025, https://www.dmgh.de/.

A substantial part of these texts have already been made available as TEI editions through the Open MGH project.

‘openMGH | Mgh.De’, accessed 15 December 2025, https://www.mgh.de/de/mgh-digital/openmgh.

With the availability of some of the texts already in TEI, it is possible to test how well the VLM is able to not only transcribe the texts, but also to translate the generated transcriptions into properly marked-up TEI.

Using this benchmarking process, I am experimenting with different prompts, parameters and image preprocessing algorithms to see what generates the highest quality edition. I am running Qwen locally, which also will enable me to do some training and finetuning of the model myself. The model is run via Python scripts to batch transcribe and process different texts per MGH publication using online available meta data.

With this approach I can expand the corpus of texts above those texts available with the entire corpus of the dMHG, and plan to incorporate print editions of other texts in my analysis of European medieval conceptions of ethnic identity. This is a faster, less labour-intensive way to make digital editions out of texts, which would allow for the quick and easy inclusion of more obscure texts, especially when funds are lacking. In the context of my field, this is especially useful for identities which due to a nationalist bias have not been studied as exhaustively, such as Aquitanian or Swabian. It could also be used to improve accessibility of texts by people belonging to oppressed or marginalised groups.

Due to the general nature of my intervention improving OCR, it could also serve as a benchmark and proof of concept for languages and texts which do not have this data available as a check. Showing the effectiveness for large scale OCR and data modelling of texts for Medieval Latin, can help scholars of languages like Ottoman Turkish or Armenian for example, judge if time investment to finetune the models is worthwhile.

References
  1. Braudel, Fernand. (1966). La Méditerranée et le monde méditerranéen à l’époque de Philippe II, Tome 1. Paris: Armand Colin.
  2. Corpus Corporum. (n.d.). Available from: https://mlat.uzh.ch/ [15 .December. 2025].
  3. Flierman, Robert. (2017). Saxon identities, AD 150-900. London: Bloomsbury Academic, an imprint of Bloomsbury Publishing Plc.
  4. Gillett, Andrew. 2006. Ethnogenesis: A Contested Model of Early Medieval Europe. History Compass. 4(2):241–260. doi.org/10.1111/J.1478-0542.2006.00311.X.
  5. He, Zhentao/ Zhang, Can/ Wu, Ziheng/ Chen, Zhenghao/ Zhan, Yufei/ Li, Yifan/ Zhang, Zhao/ Wang, Xian/ et al. (2025). doi.org/10.48550/arXiv.2506.20168.
  6. openMGH | mgh.de. (n.d.). Available from: https://www.mgh.de/de/mgh-digital/openmgh [15 .December. 2025].
  7. Pohl, Walter/ Clemens, Gantner & Payne, Richard Eds. (2012). Visions of community in the post-Roman world : the West, Byzantium and the Islamic world, 300-1100. Farnham : Ashgate,.
  8. Reimitz, Helmut. (2015). History, Frankish identity and the framing of Western ethnicity, 550-850. Cambridge: Cambridge University Press.
  9. Tang, Xiangru/ Zong, Yiming/ Phang, Jason/ Zhao, Yilun/ Zhou, Wangchunshu/ Cohan, Arman & Gerstein, Mark. (2024). doi.org/10.48550/arXiv.2309.08963.