Daejeon, July 27–31
One of the major changes brought by the digital turn in the study of written artifacts is the progressive digitization not only of their textual content but also of the text-bearing objects by their owning institutions. This large-scale availability of digital photographs favoured the development of quantitative, structured approaches described as digital paleography (Ciula 2005, Stokes 2015). Recent advances in Computer Vision using Machine Learning applied to Historical Handwriting Analysis have given rise to Computational Paleography (Stokes 2020). Interdisciplinary combinations of expertise enabled substantial progress in traditional paleographical tasks such as Writer Identification (Popović et al. 2021), Script Typology (Kestemont et al. 2017, Kurar-Barakat Berat et al. 2024) and Dating (Popović et al. 2025), although datasets made of historical documents are often small compared to the most frequently used in Computer Vision. For instance, the ICDAR 2019 HDRC-IR dataset contains 20,000 images (Christlein et al. 2019) while ImageNet has over 14 million images (11 March 2021 version, see https://www.image-net.org/). Not only is historical handwriting data available in limited quantity, but it also requires domain expertise to collect and annotate relevant datasets to address meaningful research questions. Among the theoretical questions raised by Peter Stokes ten years ago, many are still relevant today. One is “How can we use computers to help represent and investigate diachronic variation [of handwriting] as an end in itself?“ (Stokes 2015). This question is at the heart of the research project entitled ‘EGRAPSA: Retracing the evolutions of handwritings in Greco-Roman Egypt thanks to digital paleography’, funded by the Swiss National Science Foundation, whose intermediary results and future development are presented in this short paper.
Papyri preserved thanks to the dry climate of Egypt (and, in rare cases, elsewhere thanks to carbonization) offer an unrivalled source of information for ancient historians and philologists (Bagnall 2009). More than 80,000 texts in ancient Greek have been published, covering a millennium from Alexander the Great to the Arab conquest (see papyri.info). To handle such a massive amount of material, papyrologists have early engaged in digital approaches (Reggiani 2017). However, the use of Digital and Computational methods to improve paleographical studies of Greek papyri is still in its early stages. It has until now focused on precise tasks: fragment retrieval (Pirrone et al. 2021), writer identification (Christlein et al. 2022, Peer et al. 2025), character recognition (Swindall et al. 2021, Seuret et al. 2023), stylistic similarities (Marthot-Santaniello 2023) and dating (De Gregorio et al. 2024).
The goal of EGRAPSA project is to organize the newly available visual data that form papyrus images in a coherent framework and to use the new digital and computational approaches to understand what is at stake in the diversity of Greek handwritings attested on the papyri over a millennium. At the core of the project research are the questions “How much information can we infer from a handwriting?” and “With which confidence?”. An obvious parameter for making sense of the encountered diversity is the chronological evolution of handwriting that allows experts to propose a date based on paleography in the absence of other evidence. Another parameter is the type of text, since, for example, a literary piece will most often be written in a careful, set kind of script or “book hand” that differs from cursive handwriting styles encountered in legal documents or tax receipts. A third parameter is the writer: depending on the purpose, the tools, and the age, an individual’s handwriting will display variations. Thanks to a multidisciplinary team of computer scientists, papyrologists and Digital humanists, the project will explore evolutionary phenomena in a digital environment based on:
This presentation of the Egrapsa project will offer the opportunity to critically address the challenges of working with small datasets of images and computational approaches in terms of data quality, possible bias and result interpretation from the perspective of ‘transversal paleography’ (Stokes 2020) that can be used for the study of other scripts, writing material and periods. Providing new digital resources on Greek papyri and improving our understanding of the evolution of their writings will increase accessibility to this unique part of our cultural heritage for a global, non-specialist audience.