DH 2026

Daejeon, July 27–31

Thu, July 3013:40–15:10S103204-205
Short Paper

Tracing Vietnamese modernity through neologisms in colonial print (1895–1945): a computational experiment in translingual analysis

Vy Cao
Luxembourg Centre for Digital and Contemporary History, Luxembourg; Distam consortium (Digital Studies of Africa, Asia and the Middle-Est) · vy.cao@uni.lu
Frédéric Saudemont
Distam consortium (Digital Studies of Africa, Asia and the Middle-Est) · fs1510@gmail.com
  • Introduction

Since the seventeenth-century, Vietnamese romanized scripts have undergone multiple and various phases of lexical standardization and enrichment (Rageau 1984; Pham 2018). The following centuries brought significant changes as local print culture began to flourish across Southeast Asia, especially in places such as Calcutta, Bangkok, and Singapore (Rageau 1984; Pham 2018). The following centuries brought significant changes as local print culture began to flourish across Southeast Asia, especially in places such as Calcutta, Bangkok, and Singapore (Duverdier 1980; Proudfoot 1993; Ghosh 2003; D’Souza 2004). The emerging urban areas of Hanoi and Saigon in French Indochina operated the same turn by the end of the nineteenth century (Duverdier 1980; Proudfoot 1993; Ghosh 2003; D’Souza 2004). The emerging urban areas of Hanoi and Saigon in French Indochina operated the same turn by the end of the nineteenth century (V. T. Huỳnh 1971; McHale 2004; Cao 2024). In this rapidly changing environment, political and intellectual efforts sought to articulate and disseminate new vocabulary through local and transnational circuits of knowledge (V. T. Huỳnh 1971; McHale 2004; Cao 2024). In this rapidly changing environment, political and intellectual efforts sought to articulate and disseminate new vocabulary through local and transnational circuits of knowledge (Meade et al. 2024). (Meade et al. 2024).

In 1895, the first volume of the Đại Nam quốc âm tự vị – Dictionnaire de la langue annamite, compiled by Huỳnh Tịnh Của, was published in Saigon. This work is one of the earliest modern dictionaries of romanized Vietnamese (quốc ngữ) produced and edited by a Cochinchinese locutor. From 1900 onwards, hundreds of didactic and teaching books for , compiled by Huỳnh Tịnh Của, was published in Saigon. This work is one of the earliest modern dictionaries of romanized Vietnamese (quốc ngữ) produced and edited by a Cochinchinese locutor. From 1900 onwards, hundreds of didactic and teaching books for quốc ngữ circulated in Indochina, favored by the education policy implemented by the French administration. In 1917, with the support of the Directorate of Political Affairs of Indochina, Vietnamese intellectuals published Nam Phong (The Eastern Wind), a trilingual periodical which incorporated lists of Chinese, Vietnamese, and French vocabularies in its early volumes

Mostly edited by Phạm Quỳnh, a well-known scholar, Nam Phong engaged with a broad range of scientific, literary, and cultural topics.

,

Nam Phong has since been digitized and made publicly accessible through the Internet Archive, enabling large-scale computational analysis.

. These lexicons reflected a conscious effort to define vocabularies and to anchor them in a emergent public sphere (Marr 1984; McHale 2004; Peycam 2012; Nguyễn 2019).(Marr 1984; McHale 2004; Peycam 2012; Nguyễn 2019).

The computational detection of neologisms and lexical semantic change has been explored in several NLP initiatives (Schlechtweg et al. 2020). More recently, word embedding approaches have also been applied to colonial discourse in digitized historical newspapers (Schlechtweg et al. 2020). More recently, word embedding approaches have also been applied to colonial discourse in digitized historical newspapers (Zhu et al. 2025). However, such approaches remain largely untested on non-European multilingual corpora, where orthographic instability, script diversity, and the absence of annotated resources pose particular challenges.(Zhu et al. 2025). However, such approaches remain largely untested on non-European multilingual corpora, where orthographic instability, script diversity, and the absence of annotated resources pose particular challenges.

  • Research design: from embedded data to cross-lingual retrieval

This short paper introduces Rag’it, a collaborative project undertaken during a 2025 digital residency funded by the DISTAM

DIgital STudies Africa, Asia, Middle East.

consortium. The project experiment with a Retrieval Augmented Generation (RAG) pipeline to explore how usages of neologisms and their linguistic variations can be retrieved across unannotated and multilingual corpora.

The pipeline relies on a vector space model in which multilingual text segments are transformed into high-dimensional embeddings. By mapping text segments into a shared semantic space, measures such as cosine similarity allow for the querying and retrieval of segments likely to contain neologisms and their variants (Camacho-Collados and Pilehvar 2018; Liu et al. 2020). For this project, we use (Camacho-Collados and Pilehvar 2018; Liu et al. 2020). For this project, we use paraphrase-multilingual-MiniLM-L12-v2, a sentence-transformer model producing 384-dimensional dense vectors, indexed with FAISS for efficient nearest-neighbor search. This model supports Vietnamese, French, and English, and handles code-switching across mixed-language documents – a frequent feature of the colonial corpus under study.

Although embeddings can provide interesting results in semantic and cross-lingual search, it remains difficult to assess the relevance of retrieved material for specific research goals (Michail et al. 2025). (Michail et al. 2025). Rag’It intends to enhance the capability of multilingual embeddings by implementing a RAG structure using Large Language Models (LLMs), which aims to generate contextualized summaries anchored in source texts and their metadata. This structure allows researchers to evaluate retrieved segments both quantitatively and in terms of their contextual relevance to the research question.

The project is structured around two main stages. The first stage involves the optical character recognition (OCR) of the Nam Phong trilingual lexicons. This corpus offers a rich material for identifying and analyzing neologisms in the early twentieth-century Vietnamese intellectual networks. Another corpus of translated philosophical texts, published by Tân Việt editions in Hanoi in the early 1940s, has also been OCRed

This corpus includes Vietnamese translations of foundational works by popular Western thinkers such as Bergson, Kant, Nietzsche, or Freud. It has been digitized by the Bibliothèque nationale de France and is now available on Gallica.

. As a vehicle for important philosophical and conceptual vocabularies, these documents are relevant for testing and evaluating the RAG behavior and performance

This step was made possible through the use of Calfa Vision, a web-based annotation tool for both images and texts. https://vision.calfa.fr

. The second stage consists in the design of the RAG pipeline for exploratory querying across the corpus.

At its current stage, Rag’it has OCRed and published the Nam Phong trilingual lexicon dataset (Cao and Saudemont 2026). The dataset covers issues 1 to 10, 13, and 14 of (Cao and Saudemont 2026). The dataset covers issues 1 to 10, 13, and 14 of Nam Phong periodicals, representing 1,512 lexical entries. Each entry is structured across four fields corresponding to the journal’s trilingual lexicon: the romanized Vietnamese term (vi), its Han-Nôm equivalent (vi-hani), a Vietnamese-language definition (vi-def), and a French gloss (fr). Transcription was performed using a fine-tuned Vision Language Model (Qwen VL 2.5), followed by manual correction, and evaluated using KAMI (CER: 11.48%, Word Accuracy: 87.94%). This alignment across three linguistic registers constitutes the primary query interface for the RAG pipeline.

Figure Example of Nam Phong’s trilingual lexical page annotated in Calfa Vision, showing line-level segmentation and alignment of romanized Vietnamese, Chinese characters, and French glosses for OCR correction.

An experimental RAG structure has also been built within the project. The pipeline uses LangChain as the orchestration layer, coupled with locally hosted large language models (LLMs) including Mistral and PhoGPT

https://huggingface.co/vinai/models/

, to handle natural language queries in Vietnamese, French, and English. We have defined a set of multilingual queries to reflect the source materials and to test the capacity of the RAG to retrieve coherent and contextually grounded results. Our evaluation criteria focused on the extent to which the RAG output included both the text segments containing potential neologistic forms and the corresponding metadata that would allow researchers to locate and interpret these segments within their original publications.

The RAG pipeline is structured around two complementary query modes. The first operates on attested lexical forms: given a term from the structured lexical dataset in romanized Vietnamese (vi), Han-Nôm script (vi-hani), or French (fr), the system retrieves passages in which that term appears and returns it alongside its semantic context. The second operates on definitional proximity: using the Vietnamese definition field (vi-def) as a query, the system identifies passages containing terms semantically close to that definition, even when the exact form is absent. This mode is designed to trace synonyms and orthographic variants that are particularly common in early twentieth-century romanized Vietnamese.

At its current stage, the pipeline has been tested on general multilingual queries across both corpora, yielding relatively coherent and contextually grounded results. The two lexicon-based query modes described above are currently under development. Preliminary observations nonetheless highlight several challenges inherent to this line of experimentation. Ambiguities in neologistic usage and orthographic variation in historical Vietnamese texts contribute to retrieval noise and hallucination. Building a robust evaluation framework to systematically test the two lexicon-based query modes constitutes the next phase of the project.

  • Conclusion

Rag’it mobilises computational methods to address enduring questions about lexical innovation, multilingual textuality, and historical semantic change in Vietnamese print culture. The project offers a workflow for exploring neologisms in multilingual and unannotated historical corpora, which could be scaled to larger corpora of lexicons and dictionaries.

The preliminary and exploratory stage has successfully generated clean and structured dataset of the Nam Phong trilingual lexicons. It also reveals the necessity to create a coherent and robust protocol for evaluating the RAG behaviour, which could require more human and technological resources. The challenge of correctly and consistently identifying neologisms or words that are absent from the lexicon and their usage still requires more testing and evaluation. This ultimately underscores the need to create transparent documentation to allow continuous development of adaptable, context-aware models for Asian studies and the digital humanities.

References
  1. Calfa Vision. (2023). Calfa Vision: Web-based annotation tool for image and text annotation [Software]. https://calfa.fr
  2. Camacho-Collados, Jose, and Mohammad Taher Pilehvar. 2018. “From Word to Sense Embeddings: A Survey on Vector Representations of Meaning.” arXiv:1805.04032 [Cs], October 26. http://arxiv.org/abs/1805.04032.
  3. Cao, Vy. 2024. “Between the Sacred and the Secular: Publishing, Books, and Everyday Life in Colonial Cochinchina.” In Vietnam Over the Long Twentieth Century. Becoming Modern, Going Global, edited by Liam C. Kelley and Gerard Sasges. Global Vietnam: Across Time, Space and Community. Springer Singapore. https://doi.org/10.1007/978-981-97-3611-9_5.
  4. Cao, Vy, and Frédéric Saudemont. 2026. “Nam Phong Tạp Chí - Trilingual Lexicon Dataset.” Zenodo, April 6. https://doi.org/10.5281/zenodo.19438119.
  5. D’Souza, Florence. 2004. “L’émergence de l’imprimerie et de la presse écrite en Inde 1780-1830.” Bulletin des Études Indiennes (Paris) 22–23: 153–67.
  6. Duverdier, Gérald. 1980. “La transmission de l’imprimerie en Thaïlande : du catéchisme de 1796 aux impressions bouddhiques sur feuilles de latanier.” Bulletin de l’École française d’Extrême-Orient 68 (1): 209–60. https://doi.org/10.3406/befeo.1980.3331.
  7. Ghosh, Anindita. 2003. “An Uncertain ‘Coming of the Book’: Early Print Cultures in Colonial India.” Book History 6: 23–55.
  8. Huỳnh, Tịnh Của. 1895. Đại Nam quốc âm tự vị (Dictionnaire de la langue annamite). 2 vols. Imprimerie de Rey et Curiol.
  9. Huỳnh, Văn Tòng. 1971. “Histoire de la presse vietnamienne des origines à 1930.” Thèse de doctorat, Université Paris-Sorbonne.
  10. Liu, Qi, Matt J. Kusner, and Phil Blunsom. 2020. “A Survey on Contextual Embeddings.” arXiv:2003.07278 [Cs], April 13, 1–13.
  11. Marr, David G. 1984. Vietnamese Tradition on Trial, 1920-1945. University of California Press.
  12. McHale, Shawn Frederick. 2004. Print and power: confucianism, communism, and buddhism in the making of modern Vietnam. University of Hawaii Press.
  13. Meade, Ruselle, Claire Shih, and Kyung Hye Kim. 2024. Routledge Handbook of East Asian Translation. Edited by Ruselle Meade, Claire Shih, and Kyung Hye Kim. Routledge. https://doi.org/10.4324/9781003251699.
  14. Michail, Andrianos, Simon Clematide, and Rico Sennrich. 2025. “Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples.” Paper presented at EMNLP 2025. Findings of the Association for Computational Linguistics: EMNLP 2025, November. https://arxiv.org/abs/2502.08638.
  15. Nguyễn, Cindy. 2019. “Reading and Misreading: The Social Life of Libraries and Colonial Control in Vietnam, 1865-1958.” Thèse de doctorat, University of California. https://escholarship.org/uc/item/42m1f8k4.
  16. Peycam, Philippe. 2012. The Birth of Vietnamese Political Journalism: Saigon, 1916-1930. Columbia University Press.
  17. Pham, Thi Kieu Ly. 2018. “La Grammatisation du vietnamien (1615-1919) : histoire des grammaires et de l’écriture romaniisée du vietnamien.” Thèse de doctorat, Université Sorbonne Paris Cité. https://theses.hal.science/tel-03510244.
  18. Proudfoot, Ian. 1993. Early Malay Printed Books. Academy of Malay Studies and the Library, University of Malaya.
  19. Rageau, Christiane. 1984. L’imprimerie au Vietnam : de l’impression xylographique traditionnelle à la révolution du Quốc Ngữ (XIII-XIXe siècle). Bibl. municipale de Bordeaux.
  20. Schlechtweg, Dominik, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky, and Nina Tahmasebi. 2020. “SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection.” In Proceedings of the Fourteenth Workshop on Semantic Evaluation, edited by Aurelie Herbelot, Xiaodan Zhu, Alexis Palmer, Nathan Schneider, Jonathan May, and Ekaterina Shutova. International Committee for Computational Linguistics. https://doi.org/10.18653/v1/2020.semeval-1.1.
  21. Zhu, Jiaqi, Teresa Paccosi, and Marieke van Erp. 2025. “Tracing Colonial Discourse in Dutch Historical Newspapers.” In Computational Humanities Research 2025, edited by Taylor Arnold, Margherita Fantoli, and Ruben Ros. Anthology of Computers and the Humanities. https://doi.org/10.63744/SwkybkCCvmsj.