Daejeon, July 27–31
The workshop focuses on the semi-automatic collation of different textual witnesses in the context of digital editions.Therefore, the collation tool LERA (Pöckelmann et al. 2022) will be demonstrated and can be tried out in a guided hands-on session. Starting from transcribed text witnesses (XML or simple text files), the session will illustrate step by step the workflow leading to a synoptic comparison. The objective is to show how this process can support answering research questions or producing partsof a printed, digital or hybrid scholarly edition.
LERA is an interactive, web-based tool for the analysis of similarities and differences between multiple witnesses of a text. But of course it’s not the first. The analysis of text variants is an essential step in scholarly editing and thus, a wide range of automated approaches has been developed (Nury 2018). This dates back to the late 1940s with first optical solutions (Nury & Spadini 2020), before digital tools emerged, such as the well-known TUSTEP (Ott 2000), Juxta (Wheeles & Kristin 2013) or CollateX (Dekker et al. 2015). LERA differs from existing approaches as it combines three key paradigms in a single tool.
In order to be able to process even large manuscripts or books, a hierarchical collation is implemented in LERA. Thus, the texts are first segmented into larger units ‒ such as sentences or paragraphs ‒ and then automatically alignedAlgorithmic details can be found in Pöckelmann et al. 2022, section 4. to produce a comparative full text synopsis. In a second collation stepAlgorithmic details can be found in Pöckelmann 2024b, section 5.4.3., variation on a word/token level is calculated and represented through colour within the synopsis as well as within a corresponding critical apparatus.
LERA was designed to be able to handle more than two text witnesses from the beginning. A distinctive feature of the system is that each witness is compared directly with every other one. There is no need to select a base text.
Interactive intervention features allow users to manually refine the automatically generated results (segmentation, segment alignment and token alignment) directly via the graphical user interface. The analysis of text variants is further supported by integrated distant reading visualisations, such as the overview and navigation bar CATview (Pöckelmann et al. 2015) and stable word clouds (Herold et al. 2019). The resulting collations can also be exported in various formats ‒ including PDF, TeX, HTML, XML, JSON ‒ for reuse in different scholarly workflows. Figure 1 gives an impression of LERAs synoptic representation with highlighted variants (with enabled filters) and CATview on top.
Fig. 1: Screenshot of a synoptic comparison generated with LERA for a French text with four witnesses. Within the aligned segments the variation on a word level is highlighted by colour. On top, an integrated distant-reading visualisation (called CATview) gives a macro overview of the collation result.
Over the years, LERA has been used in various projects for the collation of different text witnesses. We would particularly like to emphasise the diversity of languages (as well as types of text) examined. While the tool was originally developed using texts from the French Enlightenment (Bremer et al. 2015), it has since been extended to support many other languages, including some in non-Latin alphabets, such as Arabic (Gründler & Pöckelmann 2018), Chữ Nôm (Luu et al. 2022) and Hebrew (Pöckelmann 2024a).
The workshop begins with an introductory presentation on the background of LERA and a demonstration of its main functionalities using current use cases. In the main part of the workshop, participants can use the tool with their own text material. The hands-on session includes the following aspects:
Own texts (at least two witnesses) should be prepared as UTF-8 encoded plain text (TXT) or TEI-XML. More detailed information on the import format can be found at https://lera.uzi.uni-halle.de/wiki/import. Besides, sample data will be provided too.
Learning Outcomes: Confident use of LERA's functionalities and graphical user interface, as well as a precise assessment to what extent the tool can support the realisation of your own (digital) scholarly edition projects.
Target Audience: The workshop is suitable for anyone who wants to compare different witnesses/versions of a text for research purposes. No special technical knowledge is required.
Needed Material: Participants need to bring a laptop with an up-to-date browser (LERA instances will be provided and are accessible via web) as well as their own data to compare. This should be at least two witnesses of a text that are already available as transcribed files (TXT or XML, see https://lera.uzi.uni-halle.de/wiki/import for information).
About the Instructors: The workshop will be held by Marcus Pöckelmann (Freie Universität Berlin) and Janis Dähne (Martin Luther University Halle-Wittenberg). Marcus Pöckelmann is the main developer of LERA and, as a computer scientist, has been involved in digital editions as part of multiple externally funded research projects. Janis Dähne has been researching and teaching on algorithmic problems in the context of automatic text alignment for many years. Both held similar workshops before (in English and German).