DH 2026

Daejeon, July 27–31

Workshop

Hands-On Workshop: Interactive Collation of Multiple Text Witnesses with LERA

Marcus Pöckelmann
Freie Universität Berlin, Germany · marcus.poeckelmann@fu-berlin.de
Janis Dähne
Martin Luther University Halle-Wittenberg, Germany · janis.daehne@informatik.uni-halle.de

Introduction

The workshop focuses on the semi-automatic collation of different textual witnesses in the context of digital editions.Therefore, the collation tool LERA (Pöckelmann et al. 2022) will be demonstrated and can be tried out in a guided hands-on session. Starting from transcribed text witnesses (XML or simple text files), the session will illustrate step by step the workflow leading to a synoptic comparison. The objective is to show how this process can support answering research questions or producing partsof a printed, digital or hybrid scholarly edition.

About LERA

LERA is an interactive, web-based tool for the analysis of similarities and differences between multiple witnesses of a text. But of course it’s not the first. The analysis of text variants is an essential step in scholarly editing and thus, a wide range of automated approaches has been developed (Nury 2018). This dates back to the late 1940s with first optical solutions (Nury & Spadini 2020), before digital tools emerged, such as the well-known TUSTEP (Ott 2000), Juxta (Wheeles & Kristin 2013) or CollateX (Dekker et al. 2015). LERA differs from existing approaches as it combines three key paradigms in a single tool.

Hierarchical Collation

In order to be able to process even large manuscripts or books, a hierarchical collation is implemented in LERA. Thus, the texts are first segmented into larger units ‒ such as sentences or paragraphs ‒ and then automatically alignedAlgorithmic details can be found in Pöckelmann et al. 2022, section 4. to produce a comparative full text synopsis. In a second collation stepAlgorithmic details can be found in Pöckelmann 2024b, section 5.4.3., variation on a word/token level is calculated and represented through colour within the synopsis as well as within a corresponding critical apparatus.

Equal Rank Principle

LERA was designed to be able to handle more than two text witnesses from the beginning. A distinctive feature of the system is that each witness is compared directly with every other one. There is no need to select a base text.

Semi-Automatism and Interactivity

Interactive intervention features allow users to manually refine the automatically generated results (segmentation, segment alignment and token alignment) directly via the graphical user interface. The analysis of text variants is further supported by integrated distant reading visualisations, such as the overview and navigation bar CATview (Pöckelmann et al. 2015) and stable word clouds (Herold et al. 2019). The resulting collations can also be exported in various formats ‒ including PDF, TeX, HTML, XML, JSON ‒ for reuse in different scholarly workflows. Figure 1 gives an impression of LERAs synoptic representation with highlighted variants (with enabled filters) and CATview on top.

Fig. 1: Screenshot of a synoptic comparison generated with LERA for a French text with four witnesses. Within the aligned segments the variation on a word level is highlighted by colour. On top, an integrated distant-reading visualisation (called CATview) gives a macro overview of the collation result.

Over the years, LERA has been used in various projects for the collation of different text witnesses. We would particularly like to emphasise the diversity of languages (as well as types of text) examined. While the tool was originally developed using texts from the French Enlightenment (Bremer et al. 2015), it has since been extended to support many other languages, including some in non-Latin alphabets, such as Arabic (Gründler & Pöckelmann 2018), Chữ Nôm (Luu et al. 2022) and Hebrew (Pöckelmann 2024a).

Workshop-Agenda

The workshop begins with an introductory presentation on the background of LERA and a demonstration of its main functionalities using current use cases. In the main part of the workshop, participants can use the tool with their own text material. The hands-on session includes the following aspects:

  • import of the text witnesses and their preparation for collation
  • generation of a segment alignment with LERAs automatic algorithm(s)
  • generation of a token alignment and the application of different normalisation options (filters) and colour modes
  • manual intervention through the graphical user interface
  • exploration of textual variants with the interactive visualisations integrated into LERA
  • export of achieved results for various reuse scenarios

Own texts (at least two witnesses) should be prepared as UTF-8 encoded plain text (TXT) or TEI-XML. More detailed information on the import format can be found at https://lera.uzi.uni-halle.de/wiki/import. Besides, sample data will be provided too.

Learning Outcomes: Confident use of LERA's functionalities and graphical user interface, as well as a precise assessment to what extent the tool can support the realisation of your own (digital) scholarly edition projects.

Target Audience: The workshop is suitable for anyone who wants to compare different witnesses/versions of a text for research purposes. No special technical knowledge is required.

Needed Material: Participants need to bring a laptop with an up-to-date browser (LERA instances will be provided and are accessible via web) as well as their own data to compare. This should be at least two witnesses of a text that are already available as transcribed files (TXT or XML, see https://lera.uzi.uni-halle.de/wiki/import for information).

About the Instructors: The workshop will be held by Marcus Pöckelmann (Freie Universität Berlin) and Janis Dähne (Martin Luther University Halle-Wittenberg). Marcus Pöckelmann is the main developer of LERA and, as a computer scientist, has been involved in digital editions as part of multiple externally funded research projects. Janis Dähne has been researching and teaching on algorithmic problems in the context of automatic text alignment for many years. Both held similar workshops before (in English and German).

References
  1. Bremer, Thomas / Molitor, Paul / Pöckelmann, Marcus / Ritter, Jörg / Schütz, Susanne (2015): “Zum Einsatz digitaler Methoden bei der Erstellung und Nutzung genetischer Editionen gedruckter Texte mit verschiedenen Fassungen - Das Fallbeispiel der Histoire philosphique des deux Indes von Guillaume Thomas Raynal”, in: Nutt-Kofoth, Rüdiger / Plachta, Bodo / Woesler, Winfried (eds.) Editio, 2015;29(1):29-51. DOI: 10.1515/editio-2015-004.
  2. Dekker, Ronald H. / Hulle, Dirk van / Middell, Gregor / Neyt, Vincent / Zundert, Joris van (2015) “Computer-supported collation of modern manuscripts: CollateX and the Beckett Digital Manuscript Project”, in: Digital Scholarship in the Humanities 30.3(2015):452-470.
  3. Gründler, Beatrice / Pöckelmann, Marcus (2018): “Adjusting LERA For The Comparison Of Arabic Manuscripts Of Kalīla wa-Dimna”, at: 29th international annual conference of Digital Humanities (DH2018), Mexico City, 26.-29.06.2018.
  4. Herold, Elisa / Pöckelmann, Marcus / Berg, Christian / Ritter, Jörg / Hall, Mark M. (2019): “Stable Word-Clouds for Visualising Text-Changes Over Time”, in: Doucet, Antoine / Isaac, Antoine / Golub, Koraljka / Aalberg, Trond / Jatowt, Adam (eds.) Digital Libraries for Open Knowledge - Proceedings of the 23rd International Conference on Theory and Practice of Digital Libraries (TPDL2019), Oslo, 09.-12.09.2019;224-237. DOI: 10.1007/978-3-030-30760-8_20.
  5. Luu, Thi Kim Hanh / Pöckelmann, Marcus / Ritter, Jörg / Molitor, Paul (2022): “Applying LERA for collating witnesses of The Tale of Kiều, a Vietnamese poem written in Nôm script”, in: Wang, Yifan / Murase, Tomohiro / Nagasaki, Kiyonori / Sato, Yoshihiro / Seki, Shintaro (eds.) 32nd international annual conference of Digital Humanities (DH2022) - Book of Abstracts, Tokyo, 25-29.07.2022;302-305.
  6. Nury, Elisa (2018): “Automated Collation and Digital Editions – From Theory to Practice”. PhD thesis. King’s College London.
  7. Nury, Elisa / Spadini, Elena (2020): “From Giant Despair to a New Heaven: The Early years of Automatic Collation”, in: it - Information Technology 62.2(2020):61-73. DOI: 10.1515/itit-2019-0047.
  8. Ott, Wilhelm (2000): “Strategies and tools for textual scholarship: the Tübingen system of text processing programs (TUSTEP)”, in: Literary and Linguistic Computing 15.1(2000):93-108. DOI: 10.1093/llc/15.1.93.
  9. Pöckelmann, Marcus / Medek, André / Molitor, Paul / Ritter, Jörg (2015): CATview - “Supporting The Investigation Of Text Genesis Of Large Manuscripts By An Overall Interactive Visualization Tool”, at 26th international annual conference of Digital Humanities (DH2015), Sydney, 29.06.-03.07.2015.
  10. Pöckelmann, Marcus / Medek, André / Ritter, Jörg W. / Molitor, Paul (2022): “LERA - An
    interactive platform for synoptical representations of multiple text witnesses”, in: interactive platform for synoptical representations of multiple text witnesses”, in: Digital Scholarship in the Humanities 38, 1(2023):330–346. DOI: 10.1093/llc/fqac021.
  11. Pöckelmann, Marcus (2024a): “Extensions of the Digital Collation Tool LERA for the Scholarly Edition of Keter Shem Ṭov”, in: Necker, Gerold / Rebiger, Bill (eds.) Proceedings of the Conference Editing Kabbalistic Texts, Wiesbaden: Harrassowitz Verlag, Studies in Magic and Kabbalah, 2024;2:95-110. DOI: 10.13173/9783447122412.095.
  12. Marcus Pöckelmann (2024b): “Über LERA und die Realisierung wesentlicher Anforderungen an digitale Werkzeuge zur Kollationierung verschiedener Textfassungen eines Werkes”. PhD thesis. Martin-Luther-Universität Halle-Wittenberg, 2024. DOI: 10.25673/116663.
  13. Wheeles, Dana / Kristin, Jensen (2013) “Juxta Commons”, in: Proceedings of the Digital Humanities 2013. University of Nebraska-Lincoln, pp. 545–546.