DH 2026

Daejeon, July 27–31

Workshop

Crash Course in Digital Scholarly Editions – Learning What to Do and How to Do It with Open Source and Semi-Automatic Tools

Floriane Chiffoleau
ObTIC - Sorbonne Université, France · chiffoleau.floriane@gmail.com

Introduction: the Rise of Digital Scholarly Editions

In recent years, with the rise of artificial intelligence, machine learning, and the automation of digital tasks, the digital humanities have seen a proliferation of projects involving digital scholarly editions (DSEs). DSEs are defined as “the critical representation of historic documents [...] guided by a digital paradigm in theory, method, and practice” (Driscoll / Pierazzo 2016, 23–28). These projects are not only numerous but also highly diverse in terms of structure, style, subject matter, and historical period. For instance, they may include collections of manuscripts and printed documents from Gallica (Sagot et al. 2022), parliamentary debates (Bourgeois et al. 2022), Holocaust-related archives (Bénière et al. 2024), translations of fairy tales and stories, https://hca.sdu.dk/tales/index.html (Accessed March 24, 2026) or Welsh poetry from Medieval Britain. https://myrddin.cymru/who-was-merlin/ (Accessed March 24, 2026)

In recent years, the instructors have focused on creating DSEs for a range of projects and on documenting the process for researchers—both aspects having been presented at major digital humanities conferences (Chiffoleau et al. 2021; Chagué et al. 2022).

A Pipeline: From Digitization to Publication

The process of creating a digital scholarly edition involves a series of methodical stages that are essential for transforming a physical document into a digitally enriched artifact. The workflow begins with digitization, which converts the original paper document into a digital format, typically as high-resolution images. This stage is crucial, as it determines the quality of all subsequent processing. It is followed by segmentation, which consists in identifying and analyzing the structural layout of the document, such as lines, paragraphs, or zones. Transcription then enables the extraction of textual content in a machine-readable format, allowing computational access and further manipulation. Encoding subsequently organizes this content using a markup language suited to digital representation, ensuring semantic clarity and structural coherence while making explicit the logical organization of the text. Finally, the publication stage makes the digitally enhanced document accessible online to a broader audience. This step not only presents the transformed document but also highlights its added scholarly value.

Building on previous work presented at conferences (Chiffoleau / Scheithauer 2022) and workshops, https://github.com/FloChiff/AtelierObTIC-creer-une-edition-scientifique-numerique (Accessed March 24, 2026) the instructors aim to introduce these steps and guide participants either through hands-on practice or by demonstrating techniques they can apply afterward.

Making a Digital Scholarly Edition: Crash Course with Specific Tools

A wide range of tools is available for developing digital scholarly editions, yet they are often poorly documented, difficult to locate, or scattered across different platforms. Proprietary software is generally more visible, as it benefits from stronger promotion; however, such solutions are not always suitable for researchers working with limited or no funding. In response, the instructors focus on introducing open-source, FAIR Findable, Accessible, Interoperable, Reusable. software, libraries, and tools that support each stage of the workflow. Particular attention is given to accessibility and ease of use, ensuring that participants can engage with these tools either through graphical user interfaces (GUIs) or accessible notebook environments (e.g., Jupyter notebooks or Google Colaboratory), even with little or no prior programming experience.

Using image documents brought by participants, the workshop introduces two Automatic Text Recognition (ATR) tools. Tesseract, a widely used OCR engine, is designed for printed text recognition and supports over 100 languages (Smith 2007). Kraken, on the other hand, is tailored for the segmentation and recognition of typewritten and handwritten documents (Kiessling 2019), and has seen significant improvements in recent years, now offering a wide range of models https://zenodo.org/communities/ocr_models/records (Accessed March 24, 2026) adapted to different scripts and layouts. Through guided exercises, participants will gain hands-on experience in applying these tools to their own materials.

Participants will then be introduced to the Text Encoding Initiative (TEI) and its guidelines (Burnard 2014; TEI Consortium, 2026), the standard framework for representing the structural and semantic features of texts. Through hands-on exercises, they will practice encoding metadata and textual structures, with examples tailored to different document types such as poems, theater scripts, books, correspondence, and reports. Emphasis will be placed on understanding the rationale behind encoding choices and the importance of maintaining consistency in encoding practices.

Time permitting, participants may also explore techniques for information retrieval and annotation. This includes the use of spaCy (Vasiliev 2020) for tasks such as named entity recognition, part-of-speech tagging, and lemmatization, as well as NLTK (Bird et al. 2009) for statistical, linguistic, and lexical analysis. These steps provide an introduction to how computational methods can enrich and extend scholarly editions.

Finally, the workshop addresses the publication stage and the different forms it may take. Participants will learn how to transform encoded texts into HTML/CSS outputs and explore options for dissemination, ranging from simple static websites using GitHub Pages https://docs.github.com/en/pages (Accessed March 24, 2026) to more advanced publishing environments such as TEI Publisher. https://teipublisher.com/ (Accessed March 24, 2026) This final step highlights how technical workflows can be translated into accessible and usable digital resources.

Conclusion: A Quick and Easy Introduction to Digital Scholarly Editions

This workshop aims to introduce Digital Humanities practitioners, particularly beginners in DSE, to the essential knowledge required to create a digital scholarly edition. Participants will gain an understanding of each step and be encouraged to explore their own corpora. By the end of the session, they should be able either to perform text recognition and encode documents using open-source tools or to have the necessary skills and resources to do so independently, thereby contributing to the broader dissemination of cultural heritage.

Practical information

This half-day workshop is aimed at complete beginners and is open exclusively to conference attendees. To facilitate practical exercises, participation is limited to 20–30 participants. The workshop combines theoretical presentations of all DSE steps with hands-on exercises, and participants may explore further options afterward. No prior installation is required; participants only need to bring their own computers.

The workshop primarily supports printed documents in a wide variety of languages (European, Middle Eastern, and Asian). Handwritten documents are also supported, though with a more limited range of languages (European, Arabic, and Japanese).

References
  1. Bénière, Sarah / Chiffoleau, Floriane / Scheithauer, Hugo (2024). "Streamlining the Creation of Holocaust-Related Digital Editions with Automatic Tools." Paper presented at EHRI Academic Conference – Researching the Holocaust in the Digital Age (EHRI-3), Warsaw, Poland. https://inria.hal.science/hal-04594190 (Accessed March 24, 2026.)
  2. Bird, Steven / Klein, Ewan / Loper, Edward (2009). Natural Language Processing with Python. 1st ed. Sebastopol, CA: O’Reilly Media.
  3. Bourgeois, Nicolas / Lebreton, Fanny / Pellet, Aurélien / Puren, Marie / Vernu, Pierre (2022). "Le Projet AGODA. Annoter et Publier les Débats Parlementaires Français de la fin du XIXe Siècle : Défis et Solutions." Paper presented at Présentation des projets AGODA et Gallicorpora, Bibliothèque nationale de France, Paris, France. https://hal.science/hal-03762957 (Accessed March 24, 2026.)
  4. Burnard, Lou (2014). What Is the Text Encoding Initiative? Marseille: OpenEdition Press. https://doi.org/10.4000/books.oep.426 (Accessed March 24, 2026)
  5. Chagué, Alix / Scheithauer, Hugo / Terriel, Lucas / Chiffoleau, Floriane / Tadjo-Takianpi, Yves (2022). "Take a Sip of TEI and Relax: A Proposition for an End-to-End Workflow to Enrich and Publish Data Created with Automatic Text Recognition." Paper presented at Digital Humanities 2022: Responding to Asian Diversity, Tokyo, Japan. https://hal.science/hal-03739767 (Accessed March 24, 2026)
  6. Chiffoleau, Floriane / Baillot, Anne / Ovide, Manon (2021). "A TEI-Based Publication Pipeline for Historical Egodocuments: The DAHN Project." Paper presented at Next Gen TEI – TEI Conference and Members’ Meeting, Virtual, United States. https://hal.science/hal-03451421 (Accessed March 24, 2026)
  7. Chiffoleau, Floriane / Scheithauer, Hugo (2022). "From a Collection of Documents to a Published Edition: How to Use an End-to-End Publication Pipeline." Paper presented at TEI 2022 Conference, Newcastle, United Kingdom. https://hal.science/hal-03780316 (Accessed March 24, 2026)
  8. Driscoll, Matthew James / Pierazzo, Elena (Eds.). (2016). Digital Scholarly Editing: Theories and Practices. Cambridge: Open Book Publishers. https://books.openedition.org/obp/3381 (Accessed March 24, 2026)
  9. Kiessling, Benjamin (2019). Kraken – A Universal Text Recognizer for the Humanities. Dataset. https://dataverse.nl/dataset.xhtml?persistentId=doi:10.34894/Z9G2EX (Accessed March 24, 2026)
  10. Sagot, Benoît / Romary, Laurent / Bawden, Rachel / Ortiz Suárez, Pedro Javier / Christensen, Kelly / Gabay, Simon / Pinche, Ariane / Camps, Jean-Baptiste (2022). Gallic(orpor)a: Extraction, annotation et diffusion de l’information textuelle et visuelle en diachronie longue. Paper presented at DataLab de la BnF : Restitution des travaux 2022, Paris, France. https://hal.science/hal-03930542 (Accessed March 24, 2026)
  11. Smith, Ray (2007). "An Overview of the Tesseract OCR Engine." In Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), vol. 2, 629–633. https://doi.org/10.1109/ICDAR.2007.4376991 (Accessed March 24, 2026)
  12. TEI Consortium, eds. (2026). TEI P5: Guidelines for Electronic Text Encoding and Interchange (Version 4.11.0). TEI Consortium. http://www.tei-c.org/Guidelines/P5/ (Accessed March 24, 2026)
  13. Vasiliev, Yuli (2020). Natural Language Processing with Python and spaCy: A Practical Introduction. San Francisco: No Starch Press.