DH 2026

Daejeon, July 27–31

Thu, July 3015:20–16:20S004104
Long Paper

Quantifying Translator Intervention with an Interlinear Baseline

Maciej Rapacz
AGH University of Kraków, Poland · mrapacz@agh.edu.pl

Introduction

Translation Studies has produced many theories of how translation works, but they converge on one point: there is no single "correct" rendering of a source text. The Manipulation School (Hermans 1985) makes the strongest version of this claim — every translation manipulates the source for a particular audience, purpose, or cultural context. Even less programmatic accounts agree that translating obliges the translator to adapt: syntactically, lexically, semantically, stylistically. We use translator intervention for the sum of these adaptations beyond the obligatory switch of language code: the deliberate restructuring, register shifts, and audience-driven choices that turn a source into a finished literary text. Translation Studies has described this layer extensively but largely qualitatively; computational TS has not produced a reference-free quantity for it. We propose a method for representing intervention in a vector space, capturing not only its extent but also its character.

Existing computational tools answer adjacent questions. MT evaluation, from BLEU (Papineni et al. 2002) to reference-less neural metrics like COMET-QE (Rei et al. 2020), grades quality against a reference; it does not measure how far the translator has moved from the source (Fomicheva / Specia 2016). Computational stylometry, e.g. Burrows' Delta (Burrows 2002), identifies translator fingerprints (Rybicki 2012; Mohamed et al. 2023), but its distances are relative — they say that translation A is similar to translation B without measuring the shift from the source.

We propose treating interlinear translation — a strict gloss in which target words sit under their source counterparts in source order (Shuttleworth / Cowie 2014) — as the computational baseline against which intervention is measured. Following Barthes (1953) on Writing Degree Zero, we treat this form as the limit case in which only the obligatory shift between languages has happened, before stylistic, audience-driven, or fluency-driven decisions enter. The Intervention Vector is then the embedding-space displacement of a literary translation from its interlinear counterpart; we test whether magnitude tracks the extent of intervention and whether direction tracks its character.

We validate the method on the New Testament in five languages, drawing 74 hand-picked literary translations from the targum corpus (Rapacz / Smywiński-Pohl 2026) and pairing them with five interlinear baselines we scraped for this study. Code, data, and figures are available at https://github.com/mrapacz/sighum-interlinear-vector-baselines.

Related Work

Machine Translation Evaluation. Reference-based metrics penalise divergence: neural metrics such as COMET (Rei et al. 2020) improve on BLEU (Papineni et al. 2002) but still rely on references, introducing reference bias against valid creative variation (Fomicheva / Specia 2016). Multi-reference (Wu et al. 2025) and multi-agent LLM (Kim et al. 2025) evaluators stay within the evaluative paradigm; they score quality rather than quantify intervention.

Computational Stylometry. Stylometric work focuses on attribution and clustering. Burrows' Delta (Burrows 2002) uses frequent-word distributions; multi-work studies tend to detect the author rather than the translator (Rybicki 2012), while n-gram and gradient-boosting methods succeed in multi-translator scenarios (Mohamed et al. 2023). None of these measures the departure from the source, only relative or surface-level distances within the target language.

Translationese and Corpus Linguistics. Corpus-based TS isolates universal features of translated text such as simplification and explicitation (Baker 1993), often through type-token ratio, sentence length, or POS density (Volansky et al. 2015). Translationese is typically diagnosed as fluency loss (Wein / Schneider 2024). Literary intervention is the broader phenomenon we target: deliberate restructuring whose signature is not exhausted by surface-level fingerprints.

Vector Semantics in Digital Humanities. Distributional models routinely measure diachronic semantic drift (Kutuzov et al. 2018; Hamilton et al. 2016), and BERT-style embeddings are used to study narrative structure (Wegmann et al. 2022). We adopt the same family of representations, but the shift we measure is not diachronic; it is the gap between a full translation and the interlinear baseline.

Translator Intervention in a Quantitative Setting. We know of no prior work that operationalises translator intervention as an absolute geometric quantity grounded in an explicit reference-less baseline; existing measures are either qualitative or pairwise-relative.

Methodology

The Interlinear Baseline. An interlinear gloss aligns target words to source-word order without fluency edits, defining the form rigorously enough to function as a near-source reference (Shuttleworth / Cowie 2014); Benjamin (1923) earlier called it "the archetype or ideal of all translation". The aesthetic layer is suppressed, leaving the obligatory language switch alone. We scraped interlinears from five language-specific repositories — BibleHub, Editeur BPC, Altervista, Oblubienica, Bibliatodo — removing verse markers but retaining sentence punctuation and capitalisation.

The Literary Corpus. We use the targum corpus (Rapacz / Smywiński-Pohl 2026), a multilingual collection of New Testament translations, and hand-pick 74 of them — English (16), French (14), Italian (12), Polish (16), Spanish (16) — with representatives in four orientational categories: Literal, Formal, Dynamic, and Paraphrase. The labels capture the allowance for adaptation, not theological validity. We restrict the analysis to translations from 1900 onwards to compare strategies synchronically and avoid confounds from diachronic language change.

The Intervention Vector. We embed each of the 260 New Testament chapters with Qwen3-Embedding-8B (Zhang et al. 2025), separately for every literary translation and for the interlinear baseline of the same language; the model's 32k context window holds a whole chapter in one forward pass, and using one model across all five languages keeps comparisons clean — any difference we observe reflects the translations themselves, not a checkpoint swap. We work at chapter level because verse-level segmentation is difficult: editions disagree on verse boundaries and some translations use non-standard markers (e.g. "2–6a"), and pericope-level segmentation lacks a shared schema. For each translation–chapter we then compute the displacement Vintervention = Vliterary − Vinterlinear, where Vliterary and Vinterlinear are the chapter embeddings of the literary translation and its interlinear baseline. We call this displacement the Intervention Vector and ask two questions: whether its magnitude (Euclidean norm; we also report its standard deviation across chapters) recovers the extent of intervention, and whether its direction, recovered via per-language Principal Component Analysis (PCA), carries information about character.

Results

The Spectrum of Intervention. We test whether ‖Vintervention‖ recovers the theoretical degree of intervention. Statistical tests (Mann–Whitney U with Bonferroni correction) confirm that more interventionist strategies sit further from the interlinear baseline: in English and French the four labels (Literal, Formal, Dynamic, Paraphrase) recover their full ordering, while in Italian, Polish, and Spanish the (Literal, Formal) cluster lies significantly closer to the baseline than (Dynamic, Paraphrase). At the level of individual translation pairs, one of the two is significantly closer in 90% of English pairs and 48–69% in the other languages. Without strategy labels at training time, the embedding distance reproduces the formal-to-dynamic equivalence axis that TS scholars assign by hand — a category that has historically required expert annotation falls out of an unsupervised geometric measurement.

The Range of Intervention. Median chapter-level distance to the interlinear baseline varies substantially by language — English 0.31–0.48, French 0.258–0.334, Polish 0.34–0.39, Italian 0.40–0.45, Spanish 0.44–0.66. These ranges plausibly reflect typological distance from Ancient Greek: where the target language forces more structural adaptation, translations sit further from the baseline. The mean–standard-deviation correlation is strong for English (0.83) and French (0.88), weak for Polish (0.428), weakly negative for Spanish (−0.511), and absent for Italian (0.068). We cannot, from this data alone, separate typological from representation-quality effects.

The Topology of Intervention. While magnitude measures how much a translator intervenes, it does not say how. Per-language PCA on Vintervention, with the interlinear baseline at the origin, reveals language-specific topologies: Polish forms a clear orthogonal V-shape; English and French place high-magnitude translations further from the origin without a single dominant direction; Italian and Spanish show patterns that PCA does not align with magnitude. The English panel illustrates that direction carries strategy information beyond magnitude: the First Nations Version (FNV) and the Orthodox Jewish Bible (OJB) both have high magnitude, but project along orthogonal arms — one domesticating, one foreignising. Acts 17:33 is illustrative: "Thus Paul went out from the midst of them" (interlinear) becomes "So Small Man went on his way" (FNV) and "Thus did Rav Sha'ul go out from the midst of them" (OJB). Magnitude alone would treat the two as equivalent; the geometry separates them.

Conclusion

We propose the Intervention Vector — the embedding-space difference between a literary translation and an interlinear baseline — as a reference-less framework for quantifying translator intervention. Its magnitude orders translations along the literal-to-paraphrase spectrum; the ordering is statistically significant at the group level and in roughly half of pairwise comparisons (90% in English). Per-language PCA reveals that direction carries additional information, distinguishing translations of equal magnitude that pursue opposite strategies (e.g. FNV vs. OJB).

Future work will move in two directions. First, we will test additional embedding models — both larger and from different families — to determine whether the geometry we observe is method-universal or an artefact of Qwen3-Embedding-8B; the spectrum recovery and the FNV-vs-OJB orthogonality are the two patterns we expect to behave differently across models if the method is model-bound. Second, we will probe the topology more deeply, asking which structural variables — denomination, diachronic stratum, source-edition lineage (e.g. Textus Receptus vs. Nestle–Aland) — are encoded as directions in the intervention space, and whether the same direction can be recovered across languages from translations sharing a confessional or editorial ancestry.

External links

Code, data and the embeddings are available at https://github.com/mrapacz/sighum-interlinear-vector-baselines

References
  1. Baker, Mona (1993): "Corpus Linguistics and Translation Studies: Implications and Applications", in: Baker, Mona / Francis, Gill / Tognini-Bonelli, Elena (eds.): Text and Technology: In Honour of John Sinclair. Amsterdam / Philadelphia: John Benjamins 233-250.
  2. Barthes, Roland (1953/1968): Writing Degree Zero. Translated by Annette Lavers / Colin Smith. New York: Hill and Wang.
  3. Benjamin, Walter (1923/2000): "The Task of the Translator". Translated by Harry Zohn, in: Venuti, Lawrence (ed.): The Translation Studies Reader. London / New York: Routledge 15-25.
  4. Burrows, John F. (2002): "'Delta': A Measure of Stylistic Difference and a Guide to Likely Authorship", in: Literary and Linguistic Computing 17, 3: 267-287.
  5. Fomicheva, Marina / Specia, Lucia (2016): "Reference Bias in Monolingual Machine Translation Evaluation", in: Erk, Katrin / Smith, Noah A. (eds.): Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Berlin: Association for Computational Linguistics 77-82. DOI: 10.18653/v1/P16-2013.
  6. Hamilton, William L. / Leskovec, Jure / Jurafsky, Dan (2016): "Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change", in: Erk, Katrin / Smith, Noah A. (eds.): Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Berlin: Association for Computational Linguistics 1489-1501. DOI: 10.18653/v1/P16-1141.
  7. Hermans, Theo (ed.) (1985): The Manipulation of Literature: Studies in Literary Translation. London / Sydney: Croom Helm.
  8. Kim, Junghwan / Park, Kieun / Park, Sohee / Kim, Hyunggug / Suh, Bongwon (2025): "MAS-LitEval: Multi-Agent System for Literary Translation Quality Assessment". arXiv preprint arXiv:2506.14199 <https://arxiv.org/abs/2506.14199> [08.05.2026].
  9. Kutuzov, Andrey / Øvrelid, Lilja / Szymanski, Terrence / Velldal, Erik (2018): "Diachronic Word Embeddings and Semantic Shifts: A Survey", in: Bender, Emily M. / Derczynski, Leon / Isabelle, Pierre (eds.): Proceedings of the 27th International Conference on Computational Linguistics. Santa Fe, New Mexico: Association for Computational Linguistics 1384-1397.
  10. Mohamed, Emad / Sarwar, Raheem / Mostafa, Sayed (2023): "Translator Attribution for Arabic Using Machine Learning", in: Digital Scholarship in the Humanities 38, 2: 658-666. DOI: 10.1093/llc/fqac054.
  11. Papineni, Kishore / Roukos, Salim / Ward, Todd / Zhu, Wei-Jing (2002): "BLEU: A Method for Automatic Evaluation of Machine Translation", in: Isabelle, Pierre / Charniak, Eugene / Lin, Dekang (eds.): Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. Philadelphia, Pennsylvania: Association for Computational Linguistics 311-318. DOI: 10.3115/1073083.1073135.
  12. Rapacz, Maciej / Smywiński-Pohl, Aleksander (2026): "Targum — A Multilingual New Testament Translation Corpus", in: Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). European Language Resources Association (ELRA) 7092-7105. DOI: 10.63317/2yiotxcyovir.
  13. Rei, Ricardo / Stewart, Craig / Farinha, Ana C / Lavie, Alon (2020): "COMET: A Neural Framework for MT Evaluation", in: Webber, Bonnie / Cohn, Trevor / He, Yulan / Liu, Yang (eds.): Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Online: Association for Computational Linguistics 2685-2702. DOI: 10.18653/v1/2020.emnlp-main.213.
  14. Rybicki, Jan (2012): "The Great Mystery of the (Almost) Invisible Translator: Stylometry in Translation", in: Oakes, Michael P. / Ji, Meng (eds.): Quantitative Methods in Corpus-Based Translation Studies: A Practical Guide to Descriptive Translation Research. Amsterdam / Philadelphia: John Benjamins 231-248.
  15. Shuttleworth, Mark / Cowie, Moira (2014): Dictionary of Translation Studies. London / New York: Routledge.
  16. Volansky, Vered / Ordan, Noam / Wintner, Shuly (2015): "On the Features of Translationese", in: Digital Scholarship in the Humanities 30, 1: 98-118. DOI: 10.1093/llc/fqt031.
  17. Wegmann, Anna / Schraagen, Marijn / Nguyen, Dong (2022): "Same Author or Just Same Topic? Towards Content-Independent Style Representations", in: Gella, Spandana et al. (eds.): Proceedings of the 7th Workshop on Representation Learning for NLP. Dublin: Association for Computational Linguistics 249-268. DOI: 10.18653/v1/2022.repl4nlp-1.26.
  18. Wein, Shira / Schneider, Nathan (2024): "Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation", in: Graham, Yvette / Purver, Matthew (eds.): Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). St. Julian's, Malta: Association for Computational Linguistics 753-765.
  19. Wu, Si / Wieting, John / Smith, David A. (2025): "Multiple References with Meaningful Variations Improve Literary Machine Translation". arXiv preprint arXiv:2412.18707 <https://arxiv.org/abs/2412.18707> [08.05.2026].
  20. Zhang, Yanzhao / Li, Mingxin / Long, Dingkun / Zhang, Xin / Lin, Huan / Yang, Baosong / Xie, Pengjun / Yang, An / Liu, Dayiheng / Lin, Junyang / Huang, Fei / Zhou, Jingren (2025): "Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models". arXiv preprint arXiv:2506.05176 <https://arxiv.org/abs/2506.05176> [08.05.2026].