Daejeon, July 27–31
This presentation directly responds to the DH2026 theme, “Remembering | Digital Humanities and the Memory of the World.” It reports on and examines an initiative designed to bridge the critical gap between the preservation of physical cultural heritage and digital accessibility. Our aim is to enable better interpretation of, access to, and engagement with “invaluable resources” in the East Asian world. More specifically, we explore the methodological and epistemological implications of using AI technology to reconstruct, as dynamic digital scholarly editions, local and indigenous “Memory of the World” materials that served as the basis for scholarly editions compiled one hundred years ago. We consider the significance of positioning these resources in ways that go beyond mere archival preservation, thereby fostering active research and cultural transmission within the global digital ecosystem.
The subject of this project is the “Three Editions of Buddhist Sacred Canons Preserved at Zojoji, Japan,” (UNESCO n.d.) which was inscribed on UNESCO’s Memory of the World Register in 2025. These consist of three distinct sets of woodblock-printed Tripitaka, or Buddhist canons, collected from across Japan in the early seventeenth century by Tokugawa Ieyasu, the founder of the Edo shogunate, and donated to Zojoji, the head temple of the Jodo sect. Preserved as Important Cultural Properties, these collections represent a monumental transmission of knowledge across East Asia. The three editions are as follows:
1. The Song edition, or Sixi edition, carved in twelfth-century China, consisting of 5,342 fascicles in the folded-book format, in which long sheets of paper are folded to a width of approximately 11 cm.
2. The Yuan edition, or Puning Temple edition, carved in thirteenth-century China, consisting of 5,228 fascicles in the same folded-book format.
3. The Goryeo edition, carved in thirteenth-century Korea, consisting of 1,357 volumes in bound-book format. Although the physical number of volumes differs, the substantive amount of text contained in the three editions is broadly comparable.
These three editions served as the principal source texts and collation bases for the compilation of the Taisho Shinshu Daizokyo, or Taisho Tripitaka, one hundred years ago. The Taisho Tripitaka became the international de facto standard scholarly edition for Buddhist studies and has fundamentally supported modern scholarship on East Asian Buddhism over the past century.
One hundred years after the publication of this paper-based scholarly edition, these canons have been recognized by UNESCO and digitized at the same time. Approximately 500,000 digital images have been photographed and made available on the Web, fully compatible with the International Image Interoperability Framework (IIIF) (Jodo Shu n.d.). This large-scale release of high-resolution images has had a significant impact on related research fields by democratizing access to primary sources that had previously been available only to a limited number of researchers. However, accessibility does not necessarily mean usability. Although the convenience is incomparably greater than that of paper media, the current digital archive still requires users to locate specific passages manually by visual inspection. In the age of big data, the lack of full-text searchability and granular discoverability creates inefficiencies and limits the potential for large-scale analysis and distant reading.
To address this challenge, the SAT Daizōkyō Text Database Committee (SAT Daizōkyō Text Database Committee n.d.), which has led the field of digital Buddhist studies for thirty years, has embarked on a new phase of development. We have developed a system that applies AI-OCR, or optical character recognition, to automatically transcribe the Tripitaka while simultaneously aligning the results with existing text data from the Taisho Tripitaka (Nagasaki et al. 2024). This alignment process is crucial and methodologically complex. Since the compilation of the Taisho Tripitaka one hundred years ago, scholars have identified various errors and misprints in the printed edition. Consequently, there is a dual need: to correct errors in the AI-OCR output generated from the Zojoji images, and to efficiently identify textual variants and potential editorial errors among the witnesses. Our system is designed to handle these tasks concurrently, creating a highly efficient workflow for both OCR correction and textual collation. This initiative has developed into a full-scale digital Tripitaka compilation project and has been adopted as a Grant-in-Aid for Specially Promoted Research, the highest level of research funding in Japan (Shimoda 2025).
At present, approximately 100 early-career researchers from various regions are engaged in correcting OCR texts through this system. This represents a distinctive model of “expert crowdsourcing,” in which deep domain knowledge is indispensable for deciphering variant characters and complex layouts that standard OCR cannot adequately handle. Because the corrected data is strictly linked to character images, it can be used directly for reinforcement learning of the AI-OCR engine, which uses TR-OCR. The scale of the task is enormous, requiring the correction of approximately 200 million characters in total. Nevertheless, the human-in-the-loop cycle has proven effective. After 600,000 characters had been corrected, we conducted reinforcement learning and reran the OCR process. As a result, character-level recognition accuracy improved significantly, rising from approximately 91% to 96%. This rapid improvement curve suggests that, as the project continues, the burden of manual correction will decrease and the pace of digitization will accelerate.
As a project milestone, once highly accurate transcribed texts have been completed, we plan to release them in accordance with the Text Encoding Initiative Guidelines (TEI Consortium 2025). This will not be a mere static dump of text. Rather, we aim to provide a dynamic interface through which differences among the editions can be easily confirmed. Of particular importance is that the system will allow users to instantly verify the corresponding locations on the original images. In the context of East Asian philology, the ability to confirm the shapes of characters in images of premodern books is of immense significance. Variant characters often preserve nuances or historical evidence that may be obscured by normalized Unicode text. By linking semantic text with visual evidence, we expect to greatly promote future scholarly research, enabling distant reading without sacrificing the precision of close reading.
AI technology has various aspects, both positive and negative. This project, however, demonstrates a constructive path forward. By using AI to help transform these Tripitaka, which form part of the Memory of the World, into academically reliable and highly usable resources in the digital world, we can make a significant contribution to transmitting ancient wisdom to the modern age. This initiative serves as a case study in how advanced technologies can be harmonized with traditional humanities scholarship in order to unlock the full potential of our global cultural heritage.