DH 2026

Daejeon, July 27–31

Wed, July 2914:00–15:30S076107
Short Paper

Annotating and interpreting small visual data to understand the evolution of Greek handwriting on papyri thanks to computational paleography.

Isabelle Marthot-Santaniello
University of Basel, Switzerland · i.marthot-santaniello@unibas.ch

One of the major changes brought by the digital turn in the study of written artifacts is the progressive digitization not only of their textual content but also of the text-bearing objects by their owning institutions. This large-scale availability of digital photographs favoured the development of quantitative, structured approaches described as digital paleography (Ciula 2005, Stokes 2015). Recent advances in Computer Vision using Machine Learning applied to Historical Handwriting Analysis have given rise to Computational Paleography (Stokes 2020). Interdisciplinary combinations of expertise enabled substantial progress in traditional paleographical tasks such as Writer Identification (Popović et al. 2021), Script Typology (Kestemont et al. 2017, Kurar-Barakat Berat et al. 2024) and Dating (Popović et al. 2025), although datasets made of historical documents are often small compared to the most frequently used in Computer Vision. For instance, the ICDAR 2019 HDRC-IR dataset contains 20,000 images (Christlein et al. 2019) while ImageNet has over 14 million images (11 March 2021 version, see https://www.image-net.org/). Not only is historical handwriting data available in limited quantity, but it also requires domain expertise to collect and annotate relevant datasets to address meaningful research questions. Among the theoretical questions raised by Peter Stokes ten years ago, many are still relevant today. One is “How can we use computers to help represent and investigate diachronic variation [of handwriting] as an end in itself?“ (Stokes 2015). This question is at the heart of the research project entitled ‘EGRAPSA: Retracing the evolutions of handwritings in Greco-Roman Egypt thanks to digital paleography’, funded by the Swiss National Science Foundation, whose intermediary results and future development are presented in this short paper.

Papyri preserved thanks to the dry climate of Egypt (and, in rare cases, elsewhere thanks to carbonization) offer an unrivalled source of information for ancient historians and philologists (Bagnall 2009). More than 80,000 texts in ancient Greek have been published, covering a millennium from Alexander the Great to the Arab conquest (see papyri.info). To handle such a massive amount of material, papyrologists have early engaged in digital approaches (Reggiani 2017). However, the use of Digital and Computational methods to improve paleographical studies of Greek papyri is still in its early stages. It has until now focused on precise tasks: fragment retrieval (Pirrone et al. 2021), writer identification (Christlein et al. 2022, Peer et al. 2025), character recognition (Swindall et al. 2021, Seuret et al. 2023), stylistic similarities (Marthot-Santaniello 2023) and dating (De Gregorio et al. 2024).

The goal of EGRAPSA project is to organize the newly available visual data that form papyrus images in a coherent framework and to use the new digital and computational approaches to understand what is at stake in the diversity of Greek handwritings attested on the papyri over a millennium. At the core of the project research are the questions “How much information can we infer from a handwriting?” and “With which confidence?”. An obvious parameter for making sense of the encountered diversity is the chronological evolution of handwriting that allows experts to propose a date based on paleography in the absence of other evidence. Another parameter is the type of text, since, for example, a literary piece will most often be written in a careful, set kind of script or “book hand” that differs from cursive handwriting styles encountered in legal documents or tax receipts. A third parameter is the writer: depending on the purpose, the tools, and the age, an individual’s handwriting will display variations. Thanks to a multidisciplinary team of computer scientists, papyrologists and Digital humanists, the project will explore evolutionary phenomena in a digital environment based on:

  • The collection and annotation of appropriate data: the project has started to gather precisely dated texts thanks to internal evidence, with their place of production (at a regional level) and type of text. The question of representativeness and balance across periods, locations and text types has been taken into account, as well as possible copyright bias (collections with clear CC BY licences may be overrepresented). Once the images have been collected, image areas and bounding boxes around individual characters have been defined, first manually and subsequently with the assistance of AI (see below).
  • The creation of a database on Nodegoat (Bree and Kessels 2013) that will form a panorama of handwritings. Nodegoat provides embedded visualization tools for time, space, and networks. For each selected papyrus (item), its metadata is entered into the database, along with its full image and small crops containing single characters, called cliplets. For the longest papyri, there can be more than a thousand cliplets for the most frequent letters. A timeline viewer has been developed to display cliplets (or a random selection of them for the longest texts) based on selected criteria (item, time, text type, place).
  • The creation of new tools and methods to make the most out of recent AI advances: an interface called Glyfix allows the team to curate manually the outputs of the automatic detection and recognition of the Greek characters (based on Seuret et al. 2023), moving from ca. 80 to near 99% accuracy. Similarity measurements inspired by Writer Identification (Peer et al. 2025) and script characterization (Marthot-Santaniello 2023) are being investigated with the newly generated data. Lastly, an innovative aspect of the Egrapsa project is the inclusion of hypotheses about handwriting movement (‘ductus’) as an additional level of annotation on the images. The generation of this new kind of data must be accompanied by methodological considerations and critical discussion to ensure meaningful use.

This presentation of the Egrapsa project will offer the opportunity to critically address the challenges of working with small datasets of images and computational approaches in terms of data quality, possible bias and result interpretation from the perspective of ‘transversal paleography’ (Stokes 2020) that can be used for the study of other scripts, writing material and periods. Providing new digital resources on Greek papyri and improving our understanding of the evolution of their writings will increase accessibility to this unique part of our cultural heritage for a global, non-specialist audience.

References
  1. Bagnall, Roger S. (2009): The Oxford Handbook of Papyrology. Oxford: Oxford University Press.
  2. Bree, Pim van / Kessels, Geert (2013): “Nodegoat: a web-based data management, network analysis & visualisation environment”, http://nodegoat.net
  3. Christlein, Vincent / Marthot-Santaniello, Isabelle / Mayr, Martin / Nicolaou, Anguelos / Seuret, Mathias (2022): “Writer Retrieval and Writer Identification in Greek Papyri”, in: Intertwining Graphonomics with Human Movements. IGS 2022. Lecture Notes in Computer Science, vol 13424: 76-89. https://doi.org/10.1007/978-3-031-19745-1_6
  4. Christlein, Vincent / Nicolaou, Anguelos / Seuret, Mathias / Stutzmann, Dominique / Maier, Andreas K. (2019): “ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents”, in: 2019 International Conference on Document Analysis and Recognition (ICDAR): 1505–1509. https://doi.org/10.1109/ICDAR.2019.00242
  5. Ciula, Ariana (2005): “Digital Palaeography: Using the Digital Representation of Medieval Script to Support Palaeographic Analysis”, in: Digital Medievalist 1. http://doi.org/10.16995/dm.4.
  6. De Gregorio, Giuseppe / Ferretti, Lavinia / Pena, Rodrigo C.G. / Marthot-Santaniello, Isabelle / Konstantinidou, Maria / Pavlopoulos, John (2024): “A New Framework for Error Analysis in Computational Paleographic Dating of Greek Papyri”, in: Document Analysis and Recognition – ICDAR 2024 Workshops. ICDAR 2024. Lecture Notes in Computer Science, vol 14936: 102–118. https://doi.org/10.1007/978-3-031-70642-4_7
  7. Kestemont, Mike / Christlein, Vincent / Stutzmann, Dominique (2017): “Artificial Palaeography: Computational Approaches to Identifying Script Types in Medieval Manuscripts”, in: Speculum: A Journal of Medieval Studies, 92: 86–109. https://doi.org/10.1086/694112
  8. Kurar-Barakat, Berat / Vasyutinsky-Shapira, Daria / Gogawale, Sharva / Suliman, Mohammad / Dershowitz, Nachum (2024): “Computational Paleography of Medieval Hebrew Scripts”, in: CEUR Workshop Proceedings: 707-717. https://ceur-ws.org/Vol-3834/paper42.pdf
  9. Madi, Boraq / Atamni, Nour / Tsitrinovich, Vasily / Vasyutinsky-Shapira, Daria / El-Sana, Jihad / Rabaev, Irina (2024): “Automated Dating of Medieval Manuscripts with a New Dataset”, in: Document Analysis and Recognition – ICDAR 2024 Workshops. ICDAR 2024. Lecture Notes in Computer Science, vol 14936: 45–48. https://doi.org/10.1007/978-3-031-70642-4_8
  10. Marthot-Santaniello, Isabelle / Vu, M. Tu / Serbaeva, Olga / Beurton-Aimar, Marie (2023): “Stylistic Similarities in Greek Papyri Based on Letter Shapes: A Deep Learning Approach”, in: Document Analysis and Recognition – ICDAR 2023 Workshops. ICDAR 2023. Lecture Notes in Computer Science, vol 14193: 307–323. https://doi.org/10.1007/978-3-031-41498-5_22
  11. Peer, Marco / Sablatnig, Robert / Serbaeva, Olga / Marthot-Santaniello, Isabelle (2025): “KaiRacters: Character-Level-Based Writer Retrieval for Greek Papyri”, in: Pattern Recognition. ICPR 2024. Lecture Notes in Computer Science, vol 15319: 73-88. https://doi.org/10.1007/978-3-031-78495-8_573–88
  12. Pirrone, Antoine / Beurton-Aimar, Marie / Journet, Nicholas (2021): “Self-Supervised Deep Metric Learning for Ancient Papyrus Fragments Retrieval”, in: International Journal on Document Analysis and Recognition 24: 219–34. https://doi.org/10.1007/s10032-021-00369-1
  13. Popović, Mladen / Dhali, Maruf A. / Schomaker, Lambert (2021): “Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)”, in: PLoS One, 16(4):e0249769. https://doi.org/10.1371/journal.pone.0249769
  14. Popović, Mladen / Dhali, Maruf A. / Schomaker, Lambert / van der Plicht, Johannes / Lund Rasmussen, Kaare / La Nasa Jacopo / Degano,Ilaria / Colombini, Maria Perla / Eibert, Tigchelaar (2025): “Dating ancient manuscripts using radiocarbon and AI-based writing style analysis”, in: PLoS One 20(6): e0323185. https://doi.org/10.1371/journal.pone.0323185
  15. Reggiani, Nicola (2017): Digital Papyrology I: Methods, Tools and Trends.
  16. Seuret, Mathias / Marthot-Santaniello, Isabelle / White, Stephen A. / Serbaeva Saraogi, Olga/ Agolli, Selaudin / Carrière, Guillaume / Rodriguez-Salas, Dalia / Christlein, Vincent (2023): “ICDAR 2023 Competition on Detection and Recognition of Greek Letters on Papyri”, in: Document Analysis and Recognition - ICDAR 2023. ICDAR 2023. Lecture Notes in Computer Science, vol 14188: 498-507. https://doi.org/10.1007/978-3-031-41679-8_29
  17. Stokes, Peter (2015): “Digital Approaches to Paleography and Book History: Some Challenges, Present and Future”, in: Frontiers in Digital Humanities, 2:5. https://doi.org/10.3389/fdigh.2015.00005
  18. Stokes, Peter (2020): “On Digital and Computational Humanities for Manuscript Studies: Where Have we Been, Where are we Going?”, in: Manuscript Cultures 15: 37–46. <https://www.csmc.uni-hamburg.de/publications/mc/files/articles/mc15-04-stokes.pdf>
  19. Swindall, Matthew I. / Croisdale, Gregory / Chase, C. Hunter / Keener, Ben / Williams, Alex C. / Brusuelas, James (2021): “Exploring Learning Approaches for Ancient Greek Character Recognition with Citizen Science Data”, in: 2021 IEEE 17th International Conference on eScience (eScience), 128-137. https://doi.org/10.1109/eScience51609.2021.00023