Daejeon, July 27–31
This presentation reports on the organization and TEI-based encoding of source citations found in the fragments of the Original Yupian 原本玉篇, as an initial step toward the reconstruction of this largely lost sixth-century Chinese character dictionary. The project aims to integrate traditional philological methods with recent advances in large language models (LLMs), exploring new approaches to reconstructing missing textual materials.
The Yupian, compiled by Gu Yewang 顧野王 during the Liang dynasty (6th century), was widely disseminated across East Asia and profoundly influenced the compilation of subsequent Chinese character dictionaries in both China and Japan. However, the work largely disappeared in China, and only about one-eighth survives today in the form of fragmented leaves preserved exclusively in Japan. In contrast, the Tenrei Banshō Meigi篆隷万象名義—compiled in early Heian Japan under the influence of the Yupian—offers a partially simplified yet structurally informative witness to the original text. Additionally, the Song Edition Yupian大広益会玉篇, produced in the 11th century in China, represents a later reworking that both inherits and diverges from earlier lexicographic traditions. While these works share significant overlap in entry structure and lexicographic content, they also exhibit notable differences in annotation practices, citation styles, and explanatory strategies.
Research on the Original Yupian has a long tradition in China. Hu (1989) produced the influential Yupian Jiaoshi玉篇校釋, which collated the Song edition with fragments and reconstructed lost material through textual criticism. Subsequent major works, including Zang (2008), Lü (2018), and Yao (2023), have advanced the study of the Yupian through careful philological analysis. Nevertheless, these studies rely primarily on manual comparison and editorial judgment. With the rapid development of LLMs and structured digital corpora, new possibilities for computationally assisted reconstruction have emerged.
Within the HDIC (Integrated Database of Hanzi Dictionaries in Early Japan) project at Hokkaido University, in which the author has been involved, the Tenrei Banshō Meigi, the Song edition of the Yupian, and the fragments of the Original Yupian have all been fully transcribed into plain text. While these plain-text corpora represent important resources for historical lexicography, further enrichment through structural encoding and semantic annotation is necessary to enable computational analysis. The Text Encoding Initiative (TEI P5) offers a robust international standard for modeling complex textual structures in the humanities. Recent progress in Japanese and Chinese TEI communities, including the expansion of the TEI East Asian/Japanese Working Group, has made TEI increasingly applicable to classical East Asian lexicographic materials.
This study focuses on the TEI encoding of the citation structures preserved in the fragments of the Original Yupian. These citations include quotations from major classical texts such as the Analects, the Records of the Grand Historian, the Documents, and the Mao Poetry, as well as glosses and views attributed to scholars within the exegetical tradition, including Kong Anguo, Zheng Xuan, Du Zichun, Jia Kui, and Wang Bi. By employing TEI, these heterogeneous forms of citation can be modeled within a consistent structural framework, making it possible to analyze their correspondences with the comparatively abbreviated annotation system found in the Tenrei Banshō Meigi.
The broader aim is to analyze how detailed annotations in the Original Yupian correspond to the abbreviated explanations in the Tenrei Banshō Meigi, and to design an LLM-based reverse inference model that can generate plausible reconstructions of lost passages. By aligning simplified and detailed annotation layers, the model will seek to infer fuller explanatory structures from abbreviated notes, providing candidates for reconstructing missing lexicographic entries.
This presentation reports on the first phase of this project: the systematic organization of source citations in the Original Yupian fragments and their representation through TEI markup. It also outlines how the TEI-encoded structures will be integrated with LLM-based inference in the next stage of the research. Through these combined approaches, the study aims to contribute not only to the reconstruction of the Original Yupian, but also to broader methodological discussions on applying TEI and LLMs to East Asian philological and lexicographic materials.