Daejeon, July 27–31
The COMUTE project (»Collation of Multilingual Text«) develops the first comprehensive digital environment for the systematic comparison of two or more multilingual, non-literal, and structurally divergent versions of a text. In contrast to conventional alignment methodologies predicated on near-literal, sentence-level correspondence, COMUTE is architected to accommodate complex authorial revisions, adaptive translation practices, and diachronically stratified textual variants (Balck et al. 2025).
COMUTE establishes an integrated analytical environment that facilitates the comparative examination of such heterogeneous translations and adaptations, extending the capabilities of the LERA platform, which was originally conceived for collating multiple monolingual variant traditions, to multilingual corpora (Pöckelmann et al. 2022). The present contribution delineates a series of research scenarios drawn from ongoing scholarly initiatives that demonstrate the methodological affordances of the COMUTE framework.
The first German translation of Hermann Melville’s 1851 novel »Moby-Dick; or, The Whale«, produced by Wilhelm Strüver and published in 1927, excises approximately two-thirds of the source text (Vollmer 2021, p. 55). While this phenomenon is well attested in the secondary literature, no systematic inventory of translated versus omitted passages has hitherto been undertaken. The alignment software developed within COMUTE enables the visualisation of these lacunae and facilitates granular analysis through direct juxtaposition of source and target texts (Figure 1).
The first German-language translation of J. D. Salinger’s 1951 novel »The Catcher in the Rye«, produced by Irene Muehlon and published 1954 under the title »Der Mann im Roggen« by the Swiss Diana Verlag, is not documented as containing substantive textual omissions. As anticipated, the source and target texts exhibit near-complete sentence-level correspondence, with only minimal exceptions. This finding is nonetheless methodologically significant: even where contemporary translation norms presuppose textual completeness, empirical verification through manual page-by-page collation remains impracticable.
Furthermore, the Salinger case exemplifies an extended research scenario wherein specific lexical or stylistic features may be investigated, whether through manual inspection or automated extraction procedures. Muehlon’s translation is recognised for attenuating the register of Salinger’s original, notably by substituting profanities with toned-down wording or by employing typographical elision, e. g., replacing the F-word with a dash. Beyond manual analysis, such comparative investigations can be conducted computationally, for instance by querying aligned segments that have been differentially annotated by text classifiers (e. g., hate/non-hate or toxic/non-toxic).
Numerous analogous cases merit closer examination, such as the first German translation of Astrid Lindgren’s »Pippi Longstocking« from Swedish, in which the protagonist’s anti-authoritarian comportment was systematically moderated for German readerships of the late 1940s and early 1950s.
Beyond textual omission and tonal attenuation, COMUTE supports a further research paradigm: the synchronic comparison of multiple contemporary translations of canonical works. A pertinent example involves six German translations of Charlotte Brontë’s 1847 novel »Jane Eyre«, published between 1953 and 2015 (Schmidt 2026). This investigation centres on divergent translatorial strategies, and COMUTE enables the simultaneous display of all six target texts alongside the source, substantially streamlining the collation workflow (Figure 2). Just under three decades ago, a dissertation on 21 German translations of the same novel still required »cutting out and pasting many, many scraps of paper.« (Hohn 1998, p. 5; our translation)
Applying COMUTE to different editions of the German translations of Nikos Kazantzakis’ 1946 novel »Zorba the Greek« (Soethaert 2024) highlighted two further methodological aspects. First, the source materials exhibited considerable OCR noise, yet the alignment algorithm demonstrated robust performance under these conditions. Second, the tool proved particularly valuable for identifying instances of sentence merger and segmentation divergence between source and target texts.
Building on the Greek case study, which demonstrated COMUTE’s support for non-Latin scripts, we corroborated this capability by aligning Arundhati Roy’s novel »The God of Small Things«, originally written in English, with its Hindi translation. Ongoing work in this direction continues to extend our methodological reach beyond Latin-script corpora.
The alignment of translated dramatic texts with their source versions presents distinctive challenges, as the full text incorporates genre-specific structural conventions that impede sentence-level segmentation. Speaker designations preceding speech acts, for instance, introduce considerable complexity into automated segmentation routines.
The selected case study presents an additional complication. Elizabeth Inchbald’s 1798 translation of August von Kotzebue’s 1790 play »Das Kind der Liebe«, rendered as »Lovers’ Vows«, constitutes one of the more prominent adaptations of this work, owing to its prominent role in Jane Austen’s 1814 novel »Mansfield Park«. However, Inchbald not only omitted or added scenes, as she herself acknowledges in the preface, but also adjusted the proper names of several dramatis personae to English taste. Wilhelmine, for example, becomes Agatha. Consequently, the principal obstacle to alignment was not the historical orthography of the German source but rather the structural conventions of dramatic discourse. To achieve satisfactory alignment, it was necessary to disaggregate spoken text from character designations, a procedure made feasible by the semantic markup of the dramatic texts in TEI format, which permits the discrete extraction of speech acts and stage directions while excluding the problematic speaker labels.
Another application domain involves machine-generated translations, as exemplified by a project by Svenja Guhr (UC Berkeley School of Information) and Yuri Bizzoni (Aarhus University), which is currently under review. This initiative compares translations of romance novels from US English to German, contrasting the original text with human translations and contemporary machine translations generated by LLMs, such as DeepL and Gemini. In this context, COMUTE functions as a verification mechanism for the automated assessment of translational completeness and deviation from the source text, including the analysis of censorship of explicit content.
A distinctive category of multilingual textual transmission is instantiated in the works of Hannah Arendt. Arendt believed that, after the great catastrophes of the 20th century, the world could only be understood by viewing it from different perspectives and in different languages. This strongly influenced her writing: texts written in different languages are not mere translations but rather deliberate reworkings. Her major post-war monographs were initially composed in English, with subsequent German versions prepared by the author herself. These versions diverge substantially: preliminary analysis of »The Human Condition« suggests that approximately 20% constitutes literal translation, a further 20% appears exclusively in one language version, and the remaining 60% exhibits semantic correspondence without lexical equivalence (Balck et al. 2025, Figures 3 and 4).
Arendt also revised translations of her work to ensure that the meaning and nuance were preserved. For instance, the book »Eichmann in Jerusalem« was translated into German by a translator, but Hannah Arendt was so dissatisfied with the result that she rewrote entire pages. A structured comparison of the monolingual and multilingual versions of the text is essential for understanding Hannah Arendt’s work.
These cases vividly illustrate the critical importance of bidirectional alignment. Our algorithms must not only process texts sequentially from their respective incipits but also identify transposed textual blocks that have been relocated within the macro-structure.
The foregoing use cases demonstrate that our alignment algorithm performs reliably across diverse language pairs, scripts, and textual genres. Yet beyond this technical validation, a unifying scholarly finding emerges: the cases reveal a complex continuum of translatorial practice. Texts without documented omissions, such as the Salinger and Brontë translations, nonetheless exhibit systematic register modification; ostensibly »free« adaptations, such as those by Strüver and Inchbald, preserve identifiable structural correspondences; and Arendt’s self-translations exemplify how much textual transmission occupies a middle ground that seems to resist binary categorisation altogether.
COMUTE enables the empirical mapping of this continuum at scale. Where previous scholarship could only assert degrees of divergence impressionistically, our framework permits systematic quantification and visualisation, not only of where on this spectrum any given translation falls, but crucially, of how divergence manifests: whether through omission, expansion, transposition, or register shift. This reframes our contribution from the purely technical (improved alignment of multilingual variants) to the theoretically generative: COMUTE provides the methodological infrastructure for reconceptualising translation itself as a graduated, multidimensional phenomenon rather than a categorical one.
Feedback from collaborating projects continues to inform iterative software refinement and guides the prioritisation of user requirements as we proceed toward full integration with the LERA research environment. Prior to this implementation phase, we seek scholarly dialogue to optimise our deliverables for the broader Digital Humanities community.