DH 2026

Daejeon, July 27–31

Mini-conference

From Palm Leaves to Neural Networks: Fully Engaging in the Study of Buddhist Texts

Hyoung Seok Ham
Chonnam National University, Korea, Republic of (South Korea) · hyoungseok.ham@gmail.com
Patrick McAllister
Austrian Academy of Sciences, Austria · patrick.mcallister@oeaw.ac.at
Sebastian Nehrdich
Tohoku University, Japan · nehrdbsd@gmail.com
Kiyonori Nagasaki
Keio University, Japan · nagasaki@dhii.jp

Organizers: Hyoung Seok Ham and Patrick McAllister

It is a huge challenge to study even a subset of the enormous mass of pre-modern Buddhist texts, for at least two reasons: first, it is the exception rather than the rule that one language is sufficient to study any particular text; second, in view of the sheer amount of textual material, any conclusion based on just one or only a few texts must be considered only preliminary.

Despite the emergence of digital humanities, Buddhist textual scholarship has largely retained its focus on editions that meet certain criteria (e.g., diplomatic, critical, readable), as well as on studies of individual ideas and translations of individual texts. Noteworthy exceptions to this (e.g., Qian and Radich 2021, Bingenheimer, Hung, and Wiles 2011) have shown the limits of this traditional approach. A special sort of statistical method—machine learning (ML)—is making it obvious that the two approaches will not remain separate. But it is unclear if and how qualitative and quantitative research can work together.

In the half-day conference, we want to discuss how we can maximise the mutual benefits for scholarly editors of Buddhist texts and machine-learning researchers in view of current practices and likely future developments. Our mini-conference will reflect on the following questions:

  • What can scholarly editors gain from using ML technologies?
  • Which editorial activities can ML applications replace or significantly simplify without compromising reliability?
  • How will ML methods shape our reading habits?
  • What can ML researchers gain from using scholarly editions?
  • What sort of digital editions can serve as a reliable ground truth for different ML applications?
  • How do ML applications that use editions and translations build on previous scholarship, and how can they become more trustworthy?

To tackle at least some of these questions, the proposed event invites speakers with a varied background in structural digital encoding as well as speakers who are experts in ML applications for Buddhist texts in all their complexity.

The following scholars have agreed to present their works at the mini-conference.

Patrick McAllister (Austrian Academy of Sciences, https://orcid.org/0000-0001-8043-7453) will discuss Cluttered Texts and Printed Pages: McAllister will present how he uses the TEI Guidelines as an analytical tool for philological work on Buddhist texts in a way that makes them easy to use for tasks that need “decluttered” versions, and which trade-offs need to be made for that purpose. The goal is to produce TEI-encoded scholarly editions that machine-learning tasks can work with easily.

Hyoung Seok Ham (Chonnam National University, https://orcid.org/0000-0001-8150-1445) will talk about Interpretative Layers & Network Analysis: Working with the 8th-century Tattvasaṃgraha, Ham uses encoding not for philological constitution but to tag semantic data: specific mentions of rival philosophers, usage of polemical concepts, and implicit intellectual lineages. This session demonstrates how encoding these “analytical/interpretative layers” allows us to visualise the hidden intellectual networks of Indian Buddhism.

Sebastian Nehrdich (Tohoku University, https://orcid.org/0000-0001-8728-0751) talks about Multilingual Semantic Search and Retrieval-Augmented Generation for Philology: This talk will discuss how the Dharmamitra platform is designed as an environment that integrates multilingual semantic search and retrieval-augmentation technologies to support philological work on Buddhist textual material, leading to vastly improved translation and critical editing workflows especially if multilingual versions of a given text are available.

Kiyonori Nagasaki (International Institute for Digital Humanities / Keio University, Tokyo, https://orcid.org/0000-0002-5485-0567) will present the SAT Taishō Shinshū Daizōkyō Text Database, the culmination of three decades of research in digital Buddhist studies. The first part of the presentation will focus on the integration of text and images using IIIF, as well as on the implementation of a highly efficient transcription system based on ML-driven OCR. Drawing on his long-standing efforts to build an environment for applying TEI to East Asian texts, Nagasaki will then discuss how smaller projects can achieve interoperability with global research infrastructures, thereby ensuring the sustainability and reusability of digital scholarship.

To make this event even more relevant to other conference participants, we will search for up to two further presenters once this proposal has been approved. These speakers do not need to be scholars of Buddhist texts: we will reach out to researchers with experience in ML-assisted scholarly editing, as well as to researchers with ML applications in ancient and under-resourced languages.

Proposed Format (assuming the meeting begins at 13:00)

13:00 – 13:10 (10 min): Opening Remarks

Introduction: “From Manual Encoding to AI: A Methodological Spectrum in Digital Buddhist Studies”

13:10–13:40 (30 min): TEI/Philology

Patrick McAllister: Cluttered Texts and Printed Pages

13:40–14:10 (30 min): TEI/Analysis

Hyoung Seok Ham: TEI for Intellectual Networks & Historical Analysis

14:10–14:30 (20 min): Break

14:30–15:00 (30 min): SAT Database and TEI for East Asian Texts

Kiyonori Nagasaki: Towards a TEI Framework for Buddhist Texts in East Asia

15:00–15:30 (30 min): AI Expansion

Sebastian Nehrdich: Multilingual Semantic Search and Retrieval-Augmented Generation for Philology

Optional: 15:30–16:50 (3x30 min with break): Other perspectives

Up to two other invited speakers should present relevant work.

16:50–17:20 (30 min): General Roundtable Discussion

Topic: "Synthesising Philological Rigour, Analytical Breadth, and AI in Buddhist Studies"

Contact Information of the Organizers

Ham, Hyoung Seok (Chonnam National University)

Email: hyoungseok.ham@gmail.com

Phone: (+82) 10-2817-8540

Ham is an associate professor of Indian Buddhism at Chonnam National University in the Republic of Korea. He studies the philosophical and historical aspects of the Madhyamaka tradition during the 6th to 8th centuries. He is also running a digital humanities project that investigates the intellectual network embedded in the Tattvasaṃgraha, an 8th-century Buddhist work in Sanskrit, by extensively marking up its quotation and reference information.

McAllister, Patrick (Austrian Academy of Sciences)

Email: patrick.mcallister@oeaw.ac.at

Phone: (+43) 681 84422327

McAllister is a Senior Academy Scientist at the Austrian Academy of Science’s Institute for the Cultural and Intellectual History of Asia. His primary research interest is the development of Buddhist epistemological theories during the 9th to 11th centuries (primarily in the works of Prajñākaragupta, Jñānaśrīmitra, and Ratnakīrti). He is also applies methods of the digital humanities in his research. He significantly contributes to the conceptual, methodological and technical development of the following resources:http://east.uni-hd.de/EAST, a tool to collect bibliographical and prosopographical information on the South Asian and Tibetan philosophical literature dealing with logic and argumentation. https://github.com/sarit/SARIT-corpus/blob/master/README.orgSARIT, a growing and dynamically developing library of Indic texts (mainly Sanskrit) which are encoded according to thehttps://tei-c.org/TEI Guidelines. With VEGEST, McAllister is coordinating several scholarly edition projects that are located at his home institute in Vienna.

References
  1. Bingenheimer, Marcus, Jen-Jou Hung, and Simon Wiles. 2011. “Social network visualization from TEI data.” Literary and Linguistic Computing 26.3, 271-278. doi: 10.1093/llc/fqr020
  2. Ham, Hyoung Seok. 2024. “인도 논서(śāstra) 문헌군 TEI 인코딩 전략- 해석적 층위의 데이터를 중심으로” [Encoding the Interpretative Layers of Śāstra Literature with TEI Guidelines], Korean Journal of Digital Humanities, vol. 1, no. 1, 51-72. doi: 10.23287/KJDH.2024.1.1.4 (in Korean)
  3. McAllister, Patrick. 2021. “Quotes, paraphrases, and allusions: Text reuse in sanskrit commentaries and how to encode it.” Journal of the Text Encoding Initiative, no. 13. doi:10.4000/jtei.3324.
  4. ⁣McAllister, Patrick. 2025. “VEGEST: Towards fully open and reproducible TEI workflows.” June 8. doi:10.5281/zenodo.14172008.
  5. Meelen, M., Nehrdich, S., & Keutzer, K. 2024. “Breakthroughs in Tibetan NLP & Digital Humanities.” Revue d'Etudes Tibétaines. 72, 5-25.
  6. Nagasaki, Kiyonori, Toru Tomabechi, Takahiro Kato, and Masahiro Shimoda. 2023. “Building a database for Sanskrit manuscripts compliant with IIIF and TEI.” Dejitaru Akaibu Gakkaishi 7 (s2): 103–6. doi:10.24506/jsda.7.s2_s103.
  7. Nehrdich, S., Hellwig, O., & Keutzer, K. 2024. “One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks.” In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (Findings).
  8. Qian, Lin, and Michael Radich. 2021. “A Computer-assisted Analysis of Zhu Fonian’s Original Mahayana Sutras.” Buddhist Studies Review 38.2, 145-168. doi: 10.1558/bsrv.21194