DH 2026

Daejeon, July 27–31

Poster

Fostering Engagement through East Asian Text Encoding Guidelines in the Age of Generative AI

Kiyonori Nagasaki
International Institute for Digital Humanities, Japan; Keio University · nagasaki@dhii.jp
Kazuhiro Okada
Keio University · k-okada@keio.jp
Moe Takasuka
Keio University · moe.takasuka@keio.jp
Yifan Wang
International Institute for Digital Humanities, Japan · 747.neutron@gmail.com
Masahiro Shimoda
Musashino University · shimoda@l.u-tokyo.ac.jp

Context: Engagement in the Open Science Era

Aligned with the conference theme of "Engagement," this research highlights the critical necessity of fostering meaningful connections between humanities researchers and the digital infrastructures that support their work. In the era of Open Science, there is a pressing need to establish environments that facilitate the utilization and sharing of research data. While the increasing prominence of generative AI technologies invites fresh perspectives on data analysis and modeling, the efficacy of these emerging technologies relies heavily on the quality and accessibility of the underlying data. Current metadata schemas in Japan, such as JPCOAR (JPCOAR n.d.) and JDCat (Japan Society for the Promotion of Science n.d.) , have improved discoverability but often fail to support the granular machine-readability required for deep reuse in qualitative humanities research. To bridge the gap between human interpretation and technological application, this project focuses on enhancing the machine-readability of text data through the formulation of the "Guidelines for the Construction of East Asian Text Data" (MEXT Commissioned Project n.d.) .

Technical Engagement: Adapting Guidelines for East Asia

Engagement with global standards is essential for interoperability. The project utilizes the Text Encoding Initiative (TEI) Guidelines (TEI Consortium 2025), the international de facto standard for XML-based text encoding, to create a robust data exchange format. Historically, TEI development was Euro-centric, leaving East Asian text structures underdeveloped. Through the active engagement of the SAT Daizōkyō Text Database Committee (SAT Daizōkyō Text Database Committee n.d.) and the East Asian/Japanese Special Interest Group (TEI Consortium East Asian/Japanese Special Interest Group n.d.), guidelines have been formulated to address specific regional text phenomena. The technical framework involves four approaches: utilizing existing tags such as <ab>, <seg>, and <note>, refining tags with @type attributes (e.g., for various Japanese glosses) (Nagasaki et al. 2021), using generic tags for concepts lacking specifics, and employing <anchor> tags for non-hierarchical layouts like warichu (split-line notes) (Wang et al. 2021). This active technical engagement has even led to changes in the international standard, such as the addition of official tags for ruby glosses (Okada et al. 2023).

Community Engagement: Education and Ecosystems

The "Engagement" theme emphasizes vibrancy in community interactions. Recognizing that technical guidelines alone are insufficient, the project addresses the "human" side of the equation by developing a comprehensive learning ecosystem. Since formal Digital Humanities education in Japanese universities is still limited, the project fosters community engagement through seminars, "TEI by Example" style videos, and self-learning textbooks. To empower researchers to critically examine and utilize these technologies, the project provides Google Colaboratory materials that teach Python programming for TEI data, enabling visualization and analysis. Furthermore, tools like the TEI Classical Texts Viewer <https://tei.dhii.jp/teiviewer4eaj> (Nagasaki et al. 2024) have been developed to provide immediate visual feedback on encoded data, lowering the barrier to entry and encouraging wider community participation.

Practical Engagement: Standardization for Reuse

To ensure responsible applications and sustainable engagement with data construction, the guidelines adopt the "Best Practices for TEI in Libraries" (Hawkins et al. 2018) model. This defines five levels of encoding depth, ranging from Level 1 (OCR text requiring page images) to Level 5 (semantic markup based on expert interpretation). By standardizing these levels, the project enables diverse communities to contribute data that is both manageable to produce and reliable for use by future researchers and potentially by generative AI systems seeking high-quality, structured cultural data.

Conclusion

By intertwining technical rigor with community building, this initiative demonstrates a holistic approach to "Engagement." It connects East Asian humanities with global Open Science standards, ensuring that regional cultural heritage is not only preserved but actively primed for critical reflection and utilization in an increasingly AI-driven digital landscape.

References
  1. Hawkins, Kevin / Dalmau, Michelle / Mylonas, Elli / Bauman, Syd (eds.) (2018): Best Practices for TEI in Libraries: A Guide for Mass Digitization, Automated Workflows, and Promotion of Interoperability with XML Using the TEI. Version 4.0.0. TEI Consortium. Last modified September 2018 <https://tei-c.org/extra/teiinlibraries/4.0.0/bptl-driver.html>.
  2. Japan Society for the Promotion of Science (n.d.): “JDCat Metadata Schema”
  3. <https://www.jsps.go.jp/file/storage/e-di/search/JDCat_MetaData_Schema.xlsx> [14.12.2025].
  4. JPCOAR (n.d.): “JPCOAR Schema Guidelines” <https://schema.irdb.nii.ac.jp/en> [14.12.2025].
  5. MEXT Commissioned Project: Project for the Promotion of Research and Development for DX in Humanities and Social Sciences (JPMXP1624) / Keio University / Keio Museum Commons (n.d.): “Guidelines for the Construction of East Asian Text Data” <https://teikem.kemco.keio.ac.jp/eajguidelines/> [14.12.2025].
  6. Nagasaki, Kiyonori / Inui, Yoshihiko / Kikuchi, Nobuhiko / Miyagawa, So / Ogawa, Ayumi / Horii, Hiroshi / Yoshiga, Natsuko (2021): “Man’yōshū denpon kenkyū no tame no dejitaru kiban kōchiku: Hirose-bon ‘Man’yōshū’ no kōzōka to byūwa no kaihatsu” [A Digital Platform for Genealogy of Man’yōshū], in: Jōhō Shori Gakkai Kenkyū Hōkoku [IPSJ SIG Technical Reports] 2021-CH-125, 2: 1–7.
  7. Nagasaki, Kiyonori / Honma, Jun / Shimoda, Masahiro (2024): “TEI kotenseki byūwa no kaihatsu: Higashi Ajia kotenseki ōpun saiensu no jitsugen ni mukete” [Development of TEI Classical Text Viewer: Towards Realizing Open Science for East Asian Classical Texts], in: Jinmon kagaku to konpyūta shinpojiumu ronbunshū [Proceedings of the Symposium on Humanities and Computer Processing] 2024: 59–66.
  8. Okada, Kazuhiro / Nagasaki, Kiyonori / Shimoda, Masahiro (2023): “Rubi as a Text: A Note on the Ruby Gloss Encoding”, in: Journal of the Text Encoding Initiative 14 <https://doi.org/10.4000/jtei.4159> [30.03.2023].
  9. SAT Daizōkyō Text Database Committee (n.d.): “SAT Daizōkyō Text Database” <https://21dzk.l.u-tokyo.ac.jp/SAT/> [14.12.2025].
  10. TEI Consortium (2025): TEI P5: Guidelines for Electronic Text Encoding and Interchange. Version 4.10.2. Last modified September 4 <https://tei-c.org/release/doc/tei-p5-doc/en/html/index.html>.
  11. TEI Consortium East Asian/Japanese Special Interest Group (n.d.): “TEI Kenkyūkai” [TEI Research Group] <https://tei.dhii.jp/> [14.12.2025].
  12. Wang, Yifan / Watanabe, Yoichiro / Nagasaki, Kiyonori / Shimoda, Masahiro (2021): “Xu Yiqiejing Yinyi kara miru Kanbun bunken no TEI mākuappu no kadai” [Xu Yiqiejing Yinyi and Issues on TEI Markup of Chinese Literature], in: Jinbun kagaku to konpyūta shinpojiumu ronbunshū [Proceedings of the Symposium on Humanities and Computer Languages] 2021: 234–239.