DH 2026

Daejeon, July 27–31

Poster

TEI Markup of an Early Japanese Chronicle: Focusing on Speech Segments and Quotations in the Nihon Shoki

Zifan Wu
The Graduate University for Advanced Studies, Japan; National Institute for Japanese Language and Linguistics · gs2024r007@ninjal.ac.jp
Toshinobu Ogiso
National Institute for Japanese Language and Linguistics; The Graduate University for Advanced Studies, Japan · togiso@ninjal.ac.jp

Introduction

The Nihon Shoki (720 CE), the oldest imperially commissioned historical chronicle of Japan and written entirely in Classical Chinese, is among the most important historical and linguistic sources in early Japan, yet no openly available dataset represents its internal structure in a form suitable for digital analysis. Existing online editions provide only plain text, which limits the systematic study of linguistic features and the examination of intertextual relationships such as quotations.

To address this gap, the present study structures the Nihon Shoki in XML and marks up two major categories of information: speech segments and quotations from Chinese classics and Buddhist scriptures. We also provide a TEI-compatible version to support broader interoperability, while producing a XML format for domestic search environments.

Literature Review

Research on the Nihon Shoki has long highlighted its multi-layered composition. Foundational studies by Mori (1991; 1999; 2011) and Kasai (2021) argue that variations in style and annotation reflect multiple stages of compilation. Recent quantitative approaches, such as Shi & Wang (2024), further examine textual stratification using statistical features of Japanized Classical Chinese. Studies on external sources, including Kojima (1962), Endo (2015), provide essential information for identifying quotations. In digital humanities, TEI-based encoding of East Asian classical texts has advanced through projects such as Wang et al. (2021) and Nagasaki et al. (2022). Building on this scholarship, the present study tries to create a structured XML and TEI representation of the Nihon Shoki .

Data and Challenges

This project is based on a publicly available, open-access text of the Nihon Shoki. Necessary corrections were made by consulting the Nihon Koten Bungaku Taikei (日本古典文学大系)(Sakamoto et al. 2020) and the Shinpen Nihon Koten Bungaku Zenshū (新編日本古典文学全集)(Kojima et al. 1996). Although several digital versions of the Nihon Shoki are available online, they provide only plain text without structural markup suitable for computational analysis. The Nihon Shoki presents several challenges for structured encoding.

The text contains numerous marginal notes, songs written in man’yōgana, and parallel passages introduced by expressions such as issho iwaku (一書曰), all of which require careful identification and consistent annotation

The phrase ‘一書曰’ is an expression used in the ‘Age of the Gods’ section (Volume1 & 2) of the Nihon Shoki when citing a different, alternative account alongside the main text. It indicates an important source that supplements the understanding of the main text.

.

Quotations pose a further difficulty because similar phrases appear across many Chinese classics and Buddhist scriptures, making it hard to distinguish true intertextual quotations from shorter allusions or lexical notes. Establishing stable criteria for differentiating these cases was therefore essential for creating a reliable markup.

In this project, quotations are identified when a passage of at least two characters closely matches a phrase that occurs only in a limited set of Chinese classics or Buddhist scriptures. Candidate quotations are first collected based on annotations in the Shinpen Nihon Koten Bungaku Zenshū, and then manually verified. Expressions widely attested across multiple sources or appearing only as lexical explanations in modern annotations are excluded from quotation markup. When the most plausible source differs from the modern annotation, the source attribution is revised based on textual similarity and contextual analysis.

XML Markup Strategy

The project first developed a simplified XML tagset designed for domestic search environments. This version organizes the Nihon Shoki by volumes and reigns, segments the text into basic units, and marks essential categories such as speech, quotations, parallel passages, marginal notes, and songs. At present, the domestic XML dataset has been completed for all the volumes. This lightweight schema provides a consistent structural foundation that supports efficient search and verification in systems such as Himawari (Yamaguchi & Tanaka 2005).

Building on this XML framework, we then produced a TEI-compliant version to ensure broader interoperability. In the TEI files, speech segments—including imperial edicts—are encoded with <said> and assigned speaker information. This is determined through contextual interpretation of the surrounding narrative when the speaker is not explicitly stated. Personal names are marked with <persName> and linked to entries in a centralized <listPerson>. Quotations from Chinese classics and Buddhist scriptures are represented with <bibl> elements pointing to a unified <listBibl>. We use <teiHeader> to record textual sources, encoding principles, and responsibilities.

The completed TEI files are compatible with Nagasaki et al.’s (2024) TEI Viewer for East Asian texts, allowing interactive exploration of marked-up structural and intertextual features.

Figure 1. Example of the structured data in TEI format.

Figure 2. Example of the basic interface of TEIviewer4EAJ.

Conclusion and Future Work

This project demonstrates how we marked up the textual structure of the Nihon Shoki. By providing both a custom XML schema and TEI-compliant files, we aim to contribute a reusable and interoperable resource for the international digital humanities community. We plan to release the TEI-encoded dataset online once the structural markup presented in this study has been completed for all volumes.

At present, the basic data entry for the domestic Himawari version has been largely completed, while the TEI encoding remains at a pilot stage, limited to a small number of volumes. Future work will focus on refining and extending the quotation tag system in the Himawari version and progressively expanding the TEI-encoded data to cover all 30 volumes. In addition, we are considering the introduction of a time information system to support a clearer and more structured representation of the materials. As an initial step, we plan to release a usable version of the Himawari dataset within the current fiscal year. Furthermore, we will continue to refine this encoding framework and aim for it to be applied to other classical Chinese historical materials in Japan. The TEI-encoded Nihon Shoki has the potential to serve as a foundational dataset for comparative, linguistic, and intertextual research across East Asian textual traditions.

References
  1. Endo, Keita (2015): Nihonshoki no keisei to sho shiryō [The formation of the Nihon Shoki and various sources]. Hanawashobo.
  2. Kasai, Taichi (2021): Nihonshoki dankai-hen shūron buntai chūki gohō kara mita tayō-sei to tasō-sei [The Staged Compilation of the Nihon Shoki: Diversity and Multilayeredness from the Perspective of Style, Notes, and Grammar]. Kachosha.
  3. Kojima, Noriyuki (1962): Jōdai Nihon bungaku to chū kuni bungaku: Shutten-ron o chūshin to suru hikaku bungakuteki kōsatsu,-jō [Ancient Japanese Literature and Chinese Literature: A Comparative Literary Study with a Focus on Sources, Vol. 1]. Hanawashobo.
  4. Kojima, Noriyuki / Naoki, Kojiro / Nishimiya, Kazutami / Kuranaka, Susumu / Mori, Masamori (1996): Shinpen'nihonkotenbungakuzenshū. 3 [New Complete Works of Japanese Classical Literature. 3]. Shokakukan.
  5. Mori, Hiromichi (1991): Kodai on'in to Nihonshoki no seiritsu [Ancient phonology and the creation of the Nihon Shoki]. Taishukan.
  6. Mori, Hiromichi (1999): Nihonshoki no nazowotoku: Jussaku-sha wa dare ka [Unraveling the Mystery of the Nihon Shoki: Who Was Its Author?]. Chuokoron-Shinsha.
  7. Mori, Hiromichi (2011): Nihonshoki seiritsu no shinjitsu: Kakikae no shudō-sha wa dare ka [The truth behind the creation of the Nihon Shoki: Who led the rewriting?]. Chuokoron-Shinsha.
  8. Nagasaki, Kiyonori / Nakamura, Satoru / Tanaka, Maakoto / Nishikawa, Masato / Hayashi, Ryuju / Inoue, Keijun / Shimoda, Masahiro (2022): “Current Status and Issues in Utilization of Structured Text Data--Development of a Full-Text Search System for TEI-Compliant Buddhist Scriptures—" , in: Jinmoncom 2022 Proceedings , 2022, 73-78.
  9. Nagasaki, Kiyonori / Honma, Atsushi / Shimoda, Masahiro (2024): “Development of a TEI Viewer for Classical Texts: Toward the Realization of Open Science for East Asian Classical Literature” , in: Jinmoncom 2024 Proceedings , 2024, 59-66.
  10. Sakamoto, Taro / Ienaga, Saburo / Inoue, Mitsusada / Ono, Susumu (2020): Nihonkotenbungakutaikei 67, 68: Nihonshoki [Japanese Classical Literature Series 67, 68: Nihonshoki]. Iwanamishoten.
  11. Shi, Jianjun / Wang, Dazhao (2024) : “ 基于汉文体计量特征的《日本书纪》各卷分类研究 [A Classification Study of the Volumes of the Nihon Shoki Based on the Metrological Characteristics of Classical Chinese Text] , in: Journal of Japanese Language Study and Research , 4(233), 12-25. DOI:10.13508/j.cnki.jsr.2024.04.008
  12. Wang, Yifan / Watanabe, Yoichiro / Nagasaki, Kiyonori / Shimoda, Masahiro (2021): "’Xu Yiqiejing Yinyi’ and Issues on TEI Markup of Chinese Literature” , in: Jinmoncom 2021 Proceedings , 2021, 234-239.
  13. Yamaguchi, Masaya / Tanaka, Makiro (2005): “Design and Implementation of Full Text Search System for Structured Language Resources” , in: Journal of natural language processing , 12(04), 55-77.