DH 2026

Daejeon, July 27–31

Wed, July 2909:00–10:30S025101-102
Long Paper

COMUTE in Action: Usage Scenarios for Comparing Complex Multilingual Text Variants

Frank Fischer
Freie Universität Berlin, Germany · fr.fischer@fu-berlin.de
Yashee Singh
Freie Universität Berlin, Germany · yashee.singh@fu-berlin.de
Janis Dähne
Martin Luther University Halle-Wittenberg, Germany · janis.daehne@informatik.uni-halle.de
Sascha Heße
Martin Luther University Halle-Wittenberg, Germany · sascha.hesse@informatik.uni-halle.de
Paul Molitor
Martin Luther University Halle-Wittenberg, Germany · paul.molitor@informatik.uni-halle.de
Marcus Pöckelmann
Martin Luther University Halle-Wittenberg, Germany · marcus.poeckelmann@informatik.uni-halle.de
Jörg Ritter
Martin Luther University Halle-Wittenberg, Germany · joerg.ritter@informatik.uni-halle.de
Sandra Balck
Freie Universität Berlin, Germany · sandra.balck@fu-berlin.de
Brigitte Grote
Freie Universität Berlin, Germany · brigitte.grote@fu-berlin.de
Steffen Frenzel
University of Potsdam, Germany · steffen.frenzel@uni-potsdam.de
Manfred Stede
University of Potsdam, Germany · stede@uni-potsdam.de

Introduction

The COMUTE project (»Collation of Multilingual Text«) develops the first comprehensive digital environment for the systematic comparison of two or more multilingual, non-literal, and structurally divergent versions of a text. In contrast to conventional alignment methodologies predicated on near-literal, sentence-level correspondence, COMUTE is architected to accommodate complex authorial revisions, adaptive translation practices, and diachronically stratified textual variants (Balck et al. 2025).

COMUTE establishes an integrated analytical environment that facilitates the comparative examination of such heterogeneous translations and adaptations, extending the capabilities of the LERA platform, which was originally conceived for collating multiple monolingual variant traditions, to multilingual corpora (Pöckelmann et al. 2022). The present contribution delineates a series of research scenarios drawn from ongoing scholarly initiatives that demonstrate the methodological affordances of the COMUTE framework.

Scenarios

Selective translations

The first German translation of Hermann Melville’s 1851 novel »Moby-Dick; or, The Whale«, produced by Wilhelm Strüver and published in 1927, excises approximately two-thirds of the source text (Vollmer 2021, p. 55). While this phenomenon is well attested in the secondary literature, no systematic inventory of translated versus omitted passages has hitherto been undertaken. The alignment software developed within COMUTE enables the visualisation of these lacunae and facilitates granular analysis through direct juxtaposition of source and target texts (Figure 1).

Figure Alignment of the original version of »Moby Dick« and its 1927 German translation in LERA, with omissions in the target text indicated in the synopsis visualisation above.

One-to-one translations

The first German-language translation of J. D. Salinger’s 1951 novel »The Catcher in the Rye«, produced by Irene Muehlon and published 1954 under the title »Der Mann im Roggen« by the Swiss Diana Verlag, is not documented as containing substantive textual omissions. As anticipated, the source and target texts exhibit near-complete sentence-level correspondence, with only minimal exceptions. This finding is nonetheless methodologically significant: even where contemporary translation norms presuppose textual completeness, empirical verification through manual page-by-page collation remains impracticable.

Furthermore, the Salinger case exemplifies an extended research scenario wherein specific lexical or stylistic features may be investigated, whether through manual inspection or automated extraction procedures. Muehlon’s translation is recognised for attenuating the register of Salinger’s original, notably by substituting profanities with toned-down wording or by employing typographical elision, e. g., replacing the F-word with a dash. Beyond manual analysis, such comparative investigations can be conducted computationally, for instance by querying aligned segments that have been differentially annotated by text classifiers (e. g., hate/non-hate or toxic/non-toxic).

Numerous analogous cases merit closer examination, such as the first German translation of Astrid Lindgren’s »Pippi Longstocking« from Swedish, in which the protagonist’s anti-authoritarian comportment was systematically moderated for German readerships of the late 1940s and early 1950s.

Beyond textual omission and tonal attenuation, COMUTE supports a further research paradigm: the synchronic comparison of multiple contemporary translations of canonical works. A pertinent example involves six German translations of Charlotte Brontë’s 1847 novel »Jane Eyre«, published between 1953 and 2015 (Schmidt 2026). This investigation centres on divergent translatorial strategies, and COMUTE enables the simultaneous display of all six target texts alongside the source, substantially streamlining the collation workflow (Figure 2). Just under three decades ago, a dissertation on 21 German translations of the same novel still required »cutting out and pasting many, many scraps of paper.« (Hohn 1998, p. 5; our translation)

Figure Alignment of Charlotte Brontë’s »Jane Eyre« (left column) alongside six different German translations.

Applying COMUTE to different editions of the German translations of Nikos Kazantzakis’ 1946 novel »Zorba the Greek« (Soethaert 2024) highlighted two further methodological aspects. First, the source materials exhibited considerable OCR noise, yet the alignment algorithm demonstrated robust performance under these conditions. Second, the tool proved particularly valuable for identifying instances of sentence merger and segmentation divergence between source and target texts.

Building on the Greek case study, which demonstrated COMUTE’s support for non-Latin scripts, we corroborated this capability by aligning Arundhati Roy’s novel »The God of Small Things«, originally written in English, with its Hindi translation. Ongoing work in this direction continues to extend our methodological reach beyond Latin-script corpora.

Translations of dramatic texts

The alignment of translated dramatic texts with their source versions presents distinctive challenges, as the full text incorporates genre-specific structural conventions that impede sentence-level segmentation. Speaker designations preceding speech acts, for instance, introduce considerable complexity into automated segmentation routines.

The selected case study presents an additional complication. Elizabeth Inchbald’s 1798 translation of August von Kotzebue’s 1790 play »Das Kind der Liebe«, rendered as »Lovers’ Vows«, constitutes one of the more prominent adaptations of this work, owing to its prominent role in Jane Austen’s 1814 novel »Mansfield Park«. However, Inchbald not only omitted or added scenes, as she herself acknowledges in the preface, but also adjusted the proper names of several dramatis personae to English taste. Wilhelmine, for example, becomes Agatha. Consequently, the principal obstacle to alignment was not the historical orthography of the German source but rather the structural conventions of dramatic discourse. To achieve satisfactory alignment, it was necessary to disaggregate spoken text from character designations, a procedure made feasible by the semantic markup of the dramatic texts in TEI format, which permits the discrete extraction of speech acts and stage directions while excluding the problematic speaker labels.

Machine translations

Another application domain involves machine-generated translations, as exemplified by a project by Svenja Guhr (UC Berkeley School of Information) and Yuri Bizzoni (Aarhus University), which is currently under review. This initiative compares translations of romance novels from US English to German, contrasting the original text with human translations and contemporary machine translations generated by LLMs, such as DeepL and Gemini. In this context, COMUTE functions as a verification mechanism for the automated assessment of translational completeness and deviation from the source text, including the analysis of censorship of explicit content.

Authorial self-translations

A distinctive category of multilingual textual transmission is instantiated in the works of Hannah Arendt. Arendt believed that, after the great catastrophes of the 20th century, the world could only be understood by viewing it from different perspectives and in different languages. This strongly influenced her writing: texts written in different languages are not mere translations but rather deliberate reworkings. Her major post-war monographs were initially composed in English, with subsequent German versions prepared by the author herself. These versions diverge substantially: preliminary analysis of »The Human Condition« suggests that approximately 20% constitutes literal translation, a further 20% appears exclusively in one language version, and the remaining 60% exhibits semantic correspondence without lexical equivalence (Balck et al. 2025, Figures 3 and 4).

Figure Alignment of the German and English version of Arendt’s essay »The Human Condition« showing additions and omissions in the German version (bottom row).

Figure Alignment of the German and English version of Arendt’s essay »The Human Condition«.

Arendt also revised translations of her work to ensure that the meaning and nuance were preserved. For instance, the book »Eichmann in Jerusalem« was translated into German by a translator, but Hannah Arendt was so dissatisfied with the result that she rewrote entire pages. A structured comparison of the monolingual and multilingual versions of the text is essential for understanding Hannah Arendt’s work.

These cases vividly illustrate the critical importance of bidirectional alignment. Our algorithms must not only process texts sequentially from their respective incipits but also identify transposed textual blocks that have been relocated within the macro-structure.

Conclusion and Outlook

The foregoing use cases demonstrate that our alignment algorithm performs reliably across diverse language pairs, scripts, and textual genres. Yet beyond this technical validation, a unifying scholarly finding emerges: the cases reveal a complex continuum of translatorial practice. Texts without documented omissions, such as the Salinger and Brontë translations, nonetheless exhibit systematic register modification; ostensibly »free« adaptations, such as those by Strüver and Inchbald, preserve identifiable structural correspondences; and Arendt’s self-translations exemplify how much textual transmission occupies a middle ground that seems to resist binary categorisation altogether.

COMUTE enables the empirical mapping of this continuum at scale. Where previous scholarship could only assert degrees of divergence impressionistically, our framework permits systematic quantification and visualisation, not only of where on this spectrum any given translation falls, but crucially, of how divergence manifests: whether through omission, expansion, transposition, or register shift. This reframes our contribution from the purely technical (improved alignment of multilingual variants) to the theoretically generative: COMUTE provides the methodological infrastructure for reconceptualising translation itself as a graduated, multidimensional phenomenon rather than a categorical one.

Feedback from collaborating projects continues to inform iterative software refinement and guides the prioritisation of user requirements as we proceed toward full integration with the LERA research environment. Prior to this implementation phase, we seek scholarly dialogue to optimise our deliverables for the broader Digital Humanities community.

References
  1. Sandra Balck, Janis Dähne, Fabian Etling, Frank Fischer, Steffen Frenzel, Brigitte Grote, Sascha Hesse, Paul Molitor, Marcus Pöckelmann, Jörg Ritter, Yashee Singh, Manfred Stede: Collation of Multilingual Versions of a Text: Necessity, Approach, Challenges. In: DH2025: »Accessibility and Citizenship«. 14–18 July 2025. Universidade NOVA de Lisboa.
  2. Stefanie Hohn: Charlotte Brontës »Jane Eyre« in deutscher Übersetzung. Geschichte eines kulturellen Transfers. Tübingen: Narr 1998.
  3. Marcus Pöckelmann, André Medek, Jörg Ritter, Paul Molitor: LERA – An interactive platform for synoptical representations of multiple text witnesses. In: Digital Scholarship in the Humanities, 2023;38(1):330–346. Oxford University Press 2022. https://doi.org/10.1093/llc/fqac021
  4. Henrike Schmidt: That’s (Not) What She Said: Gender Bias in neueren deutschsprachigen Übersetzungen von Charlotte Brontës »Jane Eyre«. (Master thesis.) Freie Universität Berlin 2026.
  5. Jutta Seeger-Vollmer: Schwer lesbar gleich texttreu? Wissenschaftliche Translationskritik zur Moby-Dick-Übersetzung Friedhelm Rathjens. Berlin: Frank & Timme 2021.
  6. Bart Soethaert: Circulation by Translation. In: Circulation. Ed. by Florian Fuchs, Michael Gamper, Till Kadritzke, Alexandra Ksenofontova, Jutta Müller-Tamm, Jasmin Wrobel. Articulations (March 2024). https://doi.org/10.60949/TDF0-WJ28