Daejeon, July 27–31
This study presents a digital humanities approach to the comparative study of retranslation across time, developing an LLM-assisted alignment method and using Korean retranslations of Jane Eyre as a case study to generate new evidence for interpreting the historical development of translation, literature, and publishing culture. Retranslation corpora are particularly well suited to the study of subtle cultural and linguistic shifts, because they allow researchers to compare how the same source text is rendered differently across historical moments. By holding the source text relatively constant, they make it possible to trace diachronic variation through systematic differences in honorifics, lexical choices, and other translational decisions.
Retranslation is by definition relational. It presupposes the existence of prior translations (Koskinen / Paloposki 2010), and its distinctiveness emerges not only from its relation to the source text but from an intertextual network that also includes earlier renderings (Koskinen / Paloposki 2015; Hermans 2007). Yet systematic comparison of such relations across large historical corpora has long remained difficult, since the methods needed to model large-scale textual relationships have only relatively recently begun to mature (Underwood 2019).
To quantify such changes, corpus-based translation studies and stylometric research have provided valuable tools for analyzing translational features, lexical patterning, and translator style (Baker 1995; Baker 2000; Laviosa 1998; Kenny 2001), while stylometric literary studies have likewise developed methods for identifying stylistic variation through multivariate and distributional signals (Burrows 1987; Hoover 2003; Hoover 2004; Hoover 2007). Applied to retranslation corpora, these approaches have made it possible to describe historical variation in lexical choice, stylistic preference, and stylistic consistency across editions. Yet their analytical strength has remained primarily at the aggregate level, making it difficult to measure relationships among multiple retranslations systematically and at scale, especially at the level of semantically corresponding passages. Existing work on translator style, lexical patterning, and genealogical influence has therefore often relied on close reading supplemented by relatively simple quantitative measures such as type-token ratio (TTR) (Cipriani 2023: 93–123; Ding 2024; Reynolds / Vitali 2021; Zhang 2024). Without a comprehensively aligned set of semantically corresponding units, it remains hard to locate precisely where particular features operate, to examine fine-grained translational decisions such as omission, inversion, and consolidation, or to determine whether selected passages are representative of broader textual patterns.
A more promising direction is suggested by Yang et al. (2025), who compare semantic similarity through segment embeddings after exact line-level alignment across translations. Their study is particularly valuable in demonstrating this possibility in a relatively short, aphoristic text with stable unit boundaries, but alignment in literary prose presents a more difficult problem because correspondences are often less stable across longer narrative units. Even so, their work shows that complete alignment can move analysis beyond aggregate linguistic ratios and return from quantitative patterns to specific textual loci for close reading. Yet manual alignment becomes increasingly difficult to sustain at scale, while similarity-based methods face their own limitations in literary retranslation corpora, where omission, sentence fusion, reordering, and unstable sentence boundaries frequently disrupt straightforward source–target correspondence (Levchenko 2024).
Building on this rationale, this article develops an LLM-assisted analytical framework for retranslation corpora that is designed to support both close reading and quantitative comparison across a wide range of works and languages. Its core innovation is the Aligned Meaning Segments (AMS), defined as the smallest contiguous units of meaning that make it possible to align parallel translation segments within the same source segment despite variable correspondence patterns such as 1:1, 1:N, M:1, and M:N. By establishing AMS as the basic unit of analysis, the framework makes it possible to compare retranslations not only in terms of semantic similarity and linguistic features, but also in terms of genealogical relationships among editions, including the extent to which a given translation shows continuity with earlier versions. To support this process, the framework combines a voting-based LLM self-ensemble (Wang et al. 2023) with multilingual segment embeddings and stores alignment decisions in machine-readable form so that they can be inspected, audited, and revised.
The framework comprises three stages: corpus preparation, alignment curation, and comparative analysis. Corpus preparation includes punctuation standardization, annotation removal, chapter segmentation, and sentence segmentation through Stanza (Qi et al. 2020), spaCy (Montani et al. 2023), Kiwi (Lee 2024), and KSS (Ko / Park 2021). Alignment curation proceeds through six steps—voting-based LLM self-ensemble, alignment table construction, LLM-based inspection, metric evaluation, human validation, and AMS construction—using GPT-5.2 for alignment generation and Claude Opus 4.6 as the inspector. To mitigate self-preference effects (Zheng et al. 2023), the alignment model is deliberately separated from the inspection model. Comparative analysis then proceeds through segment embeddings via voyage-4-large, Influence- and Differentiation-Ratio estimation, feature-based diachronic analysis, and metric-guided close reading.
We validate this framework through the case of Jane Eyre, a novel that has a particularly dense retranslation tradition in South Korea within a broader global translation history documented by the Prismatic Jane Eyre project (Reynolds et al. 2023). Since 1957, more than thirty-five Korean editions have been produced, highlighting South Korea as a particularly important site of Jane Eyre retranslation relative to major European languages over the same period. The corpus consists of fifteen Korean retranslations spanning 1963 to 2020, providing a historically layered record of changing readerships, publishing regimes, and translational norms. For framework validation, we present a technical demonstration based on Chapters 7 and 16, selected as contrasting discourse environments: Chapter 7 is predominantly narrative, emphasizing Jane's interiority, whereas Chapter 16 contains frequent short stretches of dialogue. The selected prompt configuration—Condition 2, with a 10:30 input and M:N flexible alignment stabilized through voting-based self-ensemble—achieved an alignment accuracy of 93.39% for Chapter 7 and 94.39% for Chapter 16.
Applied to this corpus, the AMS-based framework reveals patterned relationships among retranslations across time and connects those patterns to broader shifts in Korean publishing and translational culture. The Differentiation Ratio analysis shows three recurrent publishing cycles—1963–1978, 1982–2004, and 2010–2020—each beginning with a highly differentiated new translation and followed by later editions with progressively lower Differentiation Ratios. This pattern is consistent with broader assessments of translation quality in South Korea: a 2007 evaluation project on complete Korean translations of English and American classics published between 1945 and 2005 found that only one out of ten translations could be considered reliable, while approximately four out of ten were identified as plagiarized translations (Chung 2014). The renewed expansion of the world literature series market from the late 1990s onward—including Minumsa, Daesan World Literature, Penguin Classics Korea, Open Books, and Eulyoo Publishing—was accompanied by strategies of differentiation in selection, editorial principles, design, commentary, and the promotion of new catalogues. The pattern of overheated classical retranslation that remains especially conspicuous in Korea (Park 2018) is therefore rendered measurable through the proposed metric.
Feature-based diachronic analysis further identifies the Subject Particle Ratio (i/ga) and the Topic Particle Ratio (eun/neun) as salient markers of historical variation. The relatively high proportion of eun/neun in earlier Korean translations suggests a more topic-forward narrative style, in which characters or situations are first established as discourse topics, whereas the increasing proportion of i/ga in later editions can be interpreted as indicating a shift toward a more subject-centered mode of narration (Park / Yeon 2023). This finding stands in contrast to Ding (2024), where the most influential features were cross-linguistically general measures, suggesting that the distinction between subject marking and topic marking constitutes a particularly salient dimension of variation in Korean retranslation. Metric-guided close reading further illustrates the analytical reach of the framework: tracing AMS units with the greatest dispersion in semantic similarity reveals, for example, the lexical shift from teol yangmal (털 양말) to seutaking (스타킹) in renderings of "woollen stockings," allowing researchers to identify precisely which edition first introduced a divergent rendering and how subsequent editions diverged or converged.
More broadly, the framework addresses a persistent methodological problem in retranslation research by making large-scale alignment and comparison more systematic, auditable, and reproducible. As a digital humanities practice, it demonstrates how AI assistance can be integrated into corpus construction without compromising the interpretive transparency that humanistic inquiries require. While embedding automated LLM functions within a structured, multi-step curation pipeline, our methodology preserves the role of human judgment at every stage of inspection, validation, and revision. By transforming retranslation corpora into searchable, segment-aligned data, the framework enables systematic comparative diachronic analysis across retranslations while demonstrating a practical approach to integrating LLMs into translation studies and digital humanities.