DH 2026

Daejeon, July 27–31

Wed, July 2914:00–15:30S019105
Long Paper

Structuring the Private: Local LLMs as Cross-Cultural Annotators in Early Modern Correspondence

Mathias Johansson
Lund University, Sweden · Mathias.Johansson@kultur.lu.se
Ming Gao
Lund University, Sweden · Ming.Gao@Hist.lu.se
Cecilia Lundström
Lund University, Sweden · Cecilia.Lundstrom@Hist.lu.se
Lisa Hellman
Lund University, Sweden · Lisa.Hellman@Hist.lu.se

The digital turn in historical inquiry (Graham et al. 2015) has long promised to bridge the gap between “distant reading” (Moretti 2013) of massive corpora and the “close reading” (Richards 1929) required for nuanced interpretation. However, for historians dealing with early modern private correspondence (Daybell 2012; Gao forthcoming), this promise often hits a technological wall. These documents are characterized by their “small data” (Kitchin / Lauriault 2015) nature—specialized, highly contextual, and varying in volume—and, crucially, their paleographic and linguistic diversity. When the corpus spans the divide between the Latin alphabet and the logographic or syllabic scripts of major East Asia languages (Korean, Chinese and Japanese), the “black box” of commercial, proprietary Artificial Intelligence often fails to provide the transparency or linguistic dexterity required for rigorous scholarship.

This paper reports on a series of experiments designed to structure the analysis of early modern letters (c. 1500–1850) using Local Large Language Models (LLMs) in service of exploring how people experienced separation (Daybell 2025). Our project engages directly with two key themes of the conference: Annotating, by utilizing AI to structure unstructured historical text, and Translating, by interrogating the translingual capabilities of AI across disparate scripts. We focus on a dataset covering four linguistic contexts—Sweden, England, Korea, and Japan—and five scripts: the Latin alphabet, kanji, hangul, hanja, hiragana, and katakana.

We argue that while generalist “multilingual” models often exhibit an Anglocentric bias that flattens non-Western texts, the strategic use of language-specific local models can effectively democratize the annotation process. By treating the LLM not as an oracle but as a heuristic tool—a filtering assistant—historians can scale their analysis without losing control over privacy and interpretation.

The experiments utilized a corpus of approximately 1,500 digitalized letters, representing a subset of a larger ongoing project Moved Apart . The distribution is uneven, ranging from roughly 20 letters in the Japanese corpus to over 700 in the English corpus, reflecting the realities of archival preservation and digitization.

We have five researchers in charge of different linguistic corpora of letters, each overseeing collections ranging from several hundred to several thousand items.

The material consists of private correspondence, primarily familial interactions among the well-off, transcribed via a mix of Optical Character Recognition (OCR), Handwritten Text Recognition (HTR), and manual transcription.

A central tenet of our methodology is the rejection of commercial APIs in favor of Local LLMs. This choice is driven by two factors: Privacy and Reproducibility. Private correspondence, even from the early modern period, often resides in sensitive archives or represents cultural heritage data that should not be fed into opaque commercial training sets (Shamberger 2018). Furthermore, relying on open-weight models ensures that our experiments can be reproduced (Ries et al. 2023) by other scholars without risk of model versions changing overnight. All our experiments were conducted on consumer-grade hardware (M4 MacBook Pro) using Ollama as the inference runtime, with models quantized to fit in memory.

Keyword search is an established and trusted technique for detecting and exploring phenomena in digitalized corpora. However, it only works with a clearly defined vocabulary of keywords – which even if we could define one such a list for one of the languages presupposes that we knew, more or less, exactly what we were looking for before we began. Since expressions that indicate how people experience separations could be formulated anywhere between explicit statements to indirect, veiled hints – making such a list is a near impossible task. In contrast, language models can be very adept at dealing with nuance and ambiguity. Under the theme of Annotating | Beyond Patterns, we tasked the models with structuring the unstructured text of the letters.

We adopted prompt engineering (Sahoo et al. 2024) rather than fine-tuning — whether supervised, preference-based, or through weak supervision. The reasons are both epistemological and practical. The phenomena under study — how people experienced separation — are ephemeral and resist operationalization into fixed categories: expressions range from explicit statements to indirect, culturally embedded hints, making traditional ground truth (Liang et al. 2016) not merely expensive but conceptually fraught (Rosenwein 2015). Fine-tuning of any kind presupposes a stable task definition, which we had not yet established. Practically, six weeks of technical development time and consumer-grade hardware made iterative prompting the only feasible approach for testing across four languages and multiple models.

Instead of ground truth we relied on qualitative group evaluation (Karjus 2025) sessions where domain experts (historians specializing in the respective languages) assessed model output qualitatively — rather than against a fixed benchmark — together with the entire research team. This combination of prompt-engineering and group evaluation served as a dialogue between researchers and models: Rather than assuming a fixed capability, we adjusted our prompts, and our expectations, based on model behavior — acknowledging that effective annotation is not a one-way extraction but a negotiation between the historian’s questions and the model’s limitations. The discussions and reflections from these meetings helped us decide which prompts to adjust, which to keep and which models and prompts to abandon. We tested prompts targeting multiple variables — emotional intensity rankings, relational networks, and mentions of distance — before converging through group evaluation on a terse, four-line prompt with equivalent prompts were written in Japanese and Korean for the language-specific models, with structured output enforced via schema validation rather than free-text generation.

Consistent with known limitations of multilingual models, we observed a significant divergence in performance based on script, directly addressing the theme of Translinguality in the Era of AI. Initial tests with generalist, instruction-tuned open-weight models

Gemma3 27B, Llama 3.3 70B, and Mistral Small 3.1 24B — all instruction-tuned, run via Ollama with models quantized to fit in memory.

handled English and Swedish texts competently but, as expected, struggled with Japanese and Korean. While the models could parse the tokens, their ability to “understand” and perform even simple extraction tasks degraded sharply. We observed instances of “token loops,” where the model would output repeating nonsensical character sequences when the context window (approx. 120k tokens) was not the issue, but rather the model’s inability to navigate the logographic density of kanji or the agglutinative nature of Korean.

Furthermore, we encountered distinct Anglocentric hallucinations (Joshi et al. 2020; Cao et al. 2023). When asked to extract locations from East Asian texts, generalist models would occasionally hallucinate Western capitals (e.g., identifying “London” in a Korean text where it did not exist) or attempt to translate concepts rather than analyze them. The models displayed a bias toward English reasoning; even when prompted in the source language, the internal logic appeared to map non-Western concepts onto Western structural equivalents. We pivoted to language-specific, instruction-tuned models

EXAONE 3.5 32B (LG AI Research) for Korean; Linkbricks Horizon AI Japanese Superb V3 70B for Japanese — among several tested and discarded alternatives including DeepSeek-R1-Bllossom 70B, Trillion 7B, and SEOKDONG-llama3.1.

for Japanese and Korean. This hybrid approach dramatically improved performance. It suggests that in the “contact zone” between humans and machines, a “one-size-fits-all” multilingual model is often insufficient for historical inquiry outside the Anglo-European sphere.

A recurring obstacle was the models’ tendency toward over-interpretation (cf. sycophancy in the LLM literature (Perez et al. 2023)). Standard epistolary formulas (Bergs 2004) (e.g., “Dear Sir,” or standard honorifics) were frequently interpreted as evidence of deep emotional intimacy. However, contrary to expectations, the models handled archaic spelling and early modern grammar with remarkable resilience. The noise introduced by OCR/HTR errors (Cordell 2017) or non-standardized early modern spelling did not significantly hamper the models’ ability to parse meaning. In fact, the models would occasionally silently correct these spellings in the output—a behavior that is technically a hallucination (Ji et al. 2023) but functionally useful for data structuring.

A further, rather counter-intuitive but fortuitous observation was that the models seemed to thrive on less instruction, not more. Earlier iterations used longer, more prescriptive prompts in English — including worked examples — which paradoxically encouraged the models to hallucinate finding those examples. Furthermore, having to create a long list of possible ways to express what we were looking for would, at best, bias the models towards excluding alternative, but highly relevant expressions. When we instead gave the models the terse four-line prompt described above — without definitions or examples — it gave them freedom of interpretation. We found that they rarely missed significant elements and the false positives are often very obviously wrong. This combination renders the Local LLM a highly effective filtering tool. It allows the historian to query the dataset to find the “50 letters out of 1,500” likely to contain specific emotional patterns or relational dynamics, vastly increasing the efficiency of the research process.

Our experience challenges the binary often presented in AI discourse between “Big Data” and “Small Data.” While 1,500 letters constitute “Small Data” in the context of machine learning, they represent a significant cognitive load for a single researcher. Local LLMs bridge this gap, serving as interpreters that can handle the volume of the archive while respecting the “smallness” (privacy and context) of the individual letter.

The experiments demonstrate that the utility of AI in history is not dependent on the model achieving human-level nuance. Even a “frozen,” imperfect model running on a laptop can reveal structural patterns—provided the historian treats the output with a “trust but verify” mindset. The primary risk identified is not that the model fails, but that it succeeds too well at projecting Western structural expectations onto non-Western materials. In the presentation, we will demonstrate concrete examples of how the filtering process surfaced historically significant expressions of separation across the four linguistic corpora.

This report outlines an iterative and agile, yet reproducible approach to using Local LLMs for annotating multi-script historical correspondence. We conclude that while generalist AI models struggle with the translingual demands of global history, a curated approach using language-specific local models can unlock new methods of analysis. By having the machine structure unstructured corpora, we wish to augment the historian (Shamsieva et al. 2026) by providing high-level mappings. Nevertheless, we must remain vigilant against the Anglocentric and other biases inherent in these tools. The future of AI in Digital Humanities lies not in massive, opaque commercial models, but in the localized, iterative dialogue between the scholar and the open-source models.

References
  1. Moved Apart. https://movedapart.org/ [14 December 2025].
  2. Bergs, Alexander (2004): The Familiar Letter in Early Modern English: A Pragmatic Approach. Linguistic Society of America.
  3. Cao, Yang Trista / Sotnikova, Anna / Zhao, Jieyu / Zou, Linda X. / Rudinger, Rachel / Daumé III, Hal (2023): "Multilingual large language models leak human stereotypes across language boundaries", in: arXiv preprint arXiv:2312.07141
  4. Cordell, Ryan (2017): "’Q i-jtb the Raven’: Taking Dirty OCR Seriously", in: Book History 20:
  5. Daybell, James (2012): The Material Letter in Early Modern England: Manuscript Letters and the Culture and Practices of Letter-Writing, 1512–1635. London: Palgrave Macmillan.
  6. Daybell, James (2025): "Epistolary Technologies of Separation and Archives of Emotion" in: Green, Michael / Nørgaard, Lars Cyril (eds.): Early European Research. Turnhout, Belgium: Brepols Publishers 27–47. 10.1484/M.EER-EB.5.138238.
  7. Gao, Ming (forthcoming): "Epistolary Endurance: The ‘Grammar of Separation’ in Seventeenth-Century Letters by Japanese-Born Christian Women in Batavia", in: Journal of Global History: 10.1017/S1740022826100515.
  8. Graham, Shawn / Milligan, Ian / Weingart, Scott (2015): Exploring Big Historical Data: The Historian’s Macroscope. London: Imperial College Press.
  9. Ji, Ziwei / Lee, Nayeon / Frieske, Rita / Yu, Tiezheng / Su, Dan / Xu, Yan / Ishii, Etsuko / Bang, Ye Jin / Madotto, Andrea / Fung, Pascale (2023): "Survey of hallucination in Natural Language Generation", in: ACM Computing Surveys 55 (12): Article 248. 10.1145/3571730.
  10. Joshi, Pratik / Santy, Sebastin / Budhiraja, Amar / Bali, Kalika / Choudhury, Monojit (2020): The state and fate of linguistic diversity and inclusion in the NLP world. in: Proceedings of the 58th annual meeting of the association for computational linguistics. 6282–6293.
  11. Karjus, Andres (2025): "Machine-assisted quantitizing designs: Augmenting humanities and social sciences with Artificial Intelligence", in: Humanities and Social Sciences Communications 12 (1): 1–18.
  12. Kitchin, Rob / Lauriault, Tracey P. (2015): "Small data in the era of big data", in: GeoJournal 80: 463–475. 10.1007/s10708-014-9601-7.
  13. Liang, Jennifer J. / Tsou, Ching-Huei / Devarakonda, Murthy V. (2016): Ground truth creation for complex clinical NLP tasks – an iterative vetting approach and lessons learned. in: AMIA annual symposium proceedings. 359–367.
  14. Moretti, Franco (2013): Distant Reading. Verso Books.
  15. Perez, Ethan / Ringer, Sam / Lukosiute, Kamile / Nguyen, Karina / Chen, Edwin / Heiner, Scott / Pettit, Craig / Olsson, Catherine / Kundu, Sandipan / Kadavath, Saurav / Jones, Andy / Chen, Anna / Mann, Benjamin / Israel, Brian / Seethor, Bryan / McKinnon, Cameron / Olah, Christopher / Yan, Da / Amodei, Daniela / Amodei, Dario / Drain, Dawn / Li, Dustin / Tran-Johnson, Eli / Khundadze, Guro / Kernion, Jackson / Landis, James / Kerr, Jamie / Mueller, Jared / Hyun, Jeeyoon / Landau, Joshua / Ndousse, Kamal / Goldberg, Landon / Lovitt, Liane / Lucas, Martin / Sellitto, Michael / Zhang, Miranda / Kingsland, Neerav / Elhage, Nelson / Joseph, Nicholas / Mercado, Noemi / DasSarma, Nova / Rausch, Oliver / Larson, Robin / McCandlish, Sam / Johnston, Scott / Kravec, Shauna / El Showk, Sheer / Lanham, Tamera / Telleen-Lawton, Timothy / Brown, Tom / Henighan, Tom / Hume, Tristan / Bai, Yuntao / Hatfield-Dodds, Zac / Clark, Jack / Bowman, Samuel R. / Askell, Amanda / Grosse, Roger / Hernandez, Danny / Ganguli, Deep / Hubinger, Evan / Schiefer, Nicholas / Kaplan, Jared (2023): Discovering Language Model Behaviors with Model-Written Evaluations. in: Rogers, Anna / Boyd-Graber, Jordan / Okazaki, Naoaki (eds.): Findings of the Association for Computational Linguistics: ACL 2023. Toronto, Canada: Association for Computational Linguistics. 13387–13434. https://aclanthology.org/2023.findings-acl.847/ [5 May 2026].
  16. Richards, I. A. (1929): Practical Criticism: A Study of Literary Judgment. Cambridge: Cambridge University Press.
  17. Ries, Thorsten / Van Dalen-Oskam, Karina / Offert, Fabian (2023): "Reproducibility and explainability in Digital Humanities", in: International Journal of Digital Humanities 5 (2–3): 247–251. 10.1007/s42803-023-00078-7.
  18. Rosenwein, Barbara H. (2015): Generations of Feeling: A History of Emotions, 600–1700. Cambridge: Cambridge University Press.
  19. Sahoo, Pranab / Singh, Ayush Kumar / Saha, Sriparna / Jain, Vinija / Mondal, Samrat / Chadha, Aman (2024): "A systematic survey of prompt engineering in Large Language Models: Techniques and applications", in: arXiv preprint arXiv:2402.07927:
  20. Shamberger, Katie (2018): "Breaking rules for good? How archivists manage privacy in large-scale digitisation projects", in: Archives and Records 39 (3): 10.1080/01576895.2018.1547653.
  21. Shamsieva, Iroda / Makhmutkhodjaeva, Luiza / Tulkin, Jumaev / Irmatova, Aziza (2026): "Rethinking Humanities Education: The Impact of Artificial Intelligence and Neural Networks on Pedagogical Strategies", in: 15554: 104–120. 10.1007/978-3-031-95299-9_10.