Daejeon, July 27–31
The field of premodern ethnic identities has long been dominated by nationalist narratives. Andrew Gillett, “Ethnogenesis: A Contested Model of Early Medieval Europe,” History Compass 4, no. 2 (March 1, 2006): 199. Fernand Braudel, La Méditerranée et le monde méditerranéen à l’époque de Philippe II, (Armand Colin, 1966). Robert Flierman, Saxon Identities, AD 150-900 (Bloomsbury Academic, an imprint of Bloomsbury Publishing Plc, 2017). Robert Flierman, Saxon Identities, AD 150-900 (Bloomsbury Academic, an imprint of Bloomsbury Publishing Plc, 2017); Helmut Reimitz, History, Frankish Identity and the Framing of Western Ethnicity, 550-850 (Cambridge University Press, 2015). Patrick J. Geary, The Myth of Nations: The Medieval Origins of Europe (Princeton University Press, 2002); Visions of Community in the Post-Roman World: The West, Byzantium and the Islamic World, 300-1100, ed. Walter Pohl et al. (Ashgate, 2012).
However, to truly step out of a nationalist frame and to escape the biases of a single source on ethnic identification, widening the scope is essential. Rather than focus on one people, which results in a conceptualisation of the ethnic identity of that people, or within one source, which results in a conceptualisation of ethnic identity within that source, a plurality of peoples and sources must be studied.
For this reason, I am looking at the conceptualisation of ethnic identity in European medieval historiography through distant reading methods, which allow for a wider comparison across time and place. The main issue with this approach is that currently, not all source texts exist in a digitised edition. This limits the true potential scope of the project to already published digital editions, with all the biases that come from such a selection. For example those available on Corpus Corporum: ‘Corpus Corporum’, accessed 15 December 2025, https://mlat.uzh.ch/.
In this paper, I will show a new approach create a pipeline which takes existing scans of print editions and perform OCR, layout detection and data modelling. Using QwenLM, a Vision Language Model (VLM), to create fully marked-up digital editions, encoded in TEI. LLM’s have are known to mostly excel at data conversion tasks. Xiangru Tang et al., ‘Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?’, arXiv:2309.08963, preprint, arXiv, 4 April 2024, https://doi.org/10.48550/arXiv.2309.08963. Zhentao He et al., ‘Seeing Is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models’, arXiv:2506.20168, preprint, arXiv, 22 September 2025, https://doi.org/10.48550/arXiv.2506.20168.
To show the efficacy of this usecase I am using a corpus of texts for which already some TEI editions exist: the dMHG. The entirety of the print Monumenta Historia Germania editions have been scanned and made available online. dMGH, accessed 15 December 2025, https://www.dmgh.de/. ‘openMGH | Mgh.De’, accessed 15 December 2025, https://www.mgh.de/de/mgh-digital/openmgh.
Using this benchmarking process, I am experimenting with different prompts, parameters and image preprocessing algorithms to see what generates the highest quality edition. I am running Qwen locally, which also will enable me to do some training and finetuning of the model myself. The model is run via Python scripts to batch transcribe and process different texts per MGH publication using online available meta data.
With this approach I can expand the corpus of texts above those texts available with the entire corpus of the dMHG, and plan to incorporate print editions of other texts in my analysis of European medieval conceptions of ethnic identity. This is a faster, less labour-intensive way to make digital editions out of texts, which would allow for the quick and easy inclusion of more obscure texts, especially when funds are lacking. In the context of my field, this is especially useful for identities which due to a nationalist bias have not been studied as exhaustively, such as Aquitanian or Swabian. It could also be used to improve accessibility of texts by people belonging to oppressed or marginalised groups.
Due to the general nature of my intervention improving OCR, it could also serve as a benchmark and proof of concept for languages and texts which do not have this data available as a check. Showing the effectiveness for large scale OCR and data modelling of texts for Medieval Latin, can help scholars of languages like Ottoman Turkish or Armenian for example, judge if time investment to finetune the models is worthwhile.