Daejeon, July 27–31
1.Introduction
Ancient Chinese Encyclopedias (類書) are monumental repositories of pre-modern Chinese knowledge, compiled by excerpting and categorizing classical passages into a unique, taxonomy-based knowledge system. Beyond their encyclopedic value, Leishu are indispensable for jiyi (輯佚), which is the scholarly practice of reconstructing lost texts from fragments cited in surviving works(Zhang, 2014), as they preserve quotations from many now-lost texts.
However, traditional reconstruction methods face severe limitations when applied to the vast and complex Leishu corpus: the process is manual, laborious, time-consuming, and difficult to scale, primarily due to three factors:
First, their sheer volume makes manual passage location extremely inefficient, as many span hundreds of volumes. Second, classification systems vary drastically. For instance, the Yongle Dadian(《永樂大典》) is organized by rhyme, whereas the Gujin Tushu Jicheng(《古今圖書集成》) uses a three-level system of "Collection, Canon, and Section". This variation creates semantic and structural barriers, making cross-Leishu comparison and research very challenging. Third, knowledge units across Leishu exist as isolated, unlinked fragments, even when thematically related, preventing systematic search and analysis. Together, these challenges have blocked the creation of a unified, searchable, and analyzable Leishu knowledge system.
Digital technologies offer transformative solutions, but success requires semantic integration of disparate knowledge systems, not just digitization. Accordingly, this study focuses on the following core research question: How can we leverage formal knowledge representation and intelligent systems to digitally reconstruct the fragmented knowledge within Leishu, in order to enable large-scale, efficient reconstructive compilation and support integrated scholarly analysis?
Thus, we propose an ontology approach. Ontologies provide formal, shared conceptual models that can explicitly define domain entities and relationships. Our Leishu Ontology creates a semantic framework to reconcile cross-encyclopedia heterogeneity, serveing as the backbone for an intelligent platform that aggregates content, supports advanced variant searches, and facilitates long text reconstruction and encyclopedic knowledge exploration.
2. Literature Review
2.1 Digital Approaches to Encyclopedic and Fragmentary Texts
The digital reconstruction of classical knowledge systems has become crucial in digital humanities. Digital textual criticism combines traditional philological principles with computational techniques to analyze variants and reconstruct textual histories (Pierazzo, 2015). Tools like HyperCollate enables scalable, algorithmic collation of non-linear, multi-version texts through XML graph comparison. (Bleeker et al., 2022). Complementary projects like VisColl, model the physical collation structure of manuscripts to visualize gatherings, quires, and bindings (Porter et al., 2017). These systems illustrate how digital tools integrate material and textual evidence to support variant analysis and edition building.
The emerging field of digital fragmentology focus on virtual reconstruction of dispersed fragments. Projects like Fragmentarium and the Digital Vatican Library use IIIF frameworks to reunite manuscript folia via high-resolution imaging and semantic annotation, enabling cross-institutional collaborative analysis and linked data modeling of fragment reconstruction. Projects focusing on lost text recovery or quotation networks such as Perseus Digital Library and Open Greek and Latin Project, apply citation indexing to reconstruct textual traditions (Crane et al., 2006). Computational linguistics research in text reuse detection extends this paradigm by using NLP algorithms to identify parallel passages, paraphrases, and quotations in large corpora (Mahadevan et al., 2025). Together, these provide a methodological foundation for automated jiyi (輯佚) assistance.
2.2 Knowledge Organization for Classical Texts
Advances in knowledge organization enable cultural heritage data to be modeled as interoperable knowledge graphs. General frameworks like CIDOC CRM and FRBRoo provide theoretical paradigms for ancient text ontology development (Melo 2023). Large-scale infrastructures such as Europeana employ these standards to facilitate interoperability across cultural repositories (Charles et al., 2020).
In Sinological digital humanities, ontology have proven effective for representing Chinese textual and bibliographic knowledge. Wang et al. (2023) developed the CABC Ontology for ancient Chinese book catalogs within the FRBR/BIBFRAME paradigm, and Wang et al. (2024) created the TCT Ontology to formalize temporal expressions in Chinese historiographical texts. Cai (2023) applied CIDOC-CRM to traditional Chinese papermaking, and Zheng et al. (2024) constructed a knowledge graph for ancient scientific literature using BERT-based named entity recognition. These studies demonstrate the adaptability of ontology frameworks to complex classical data.
Despite these achievements, no existing ontology captures the distinctive structural and epistemological features of ancient chinese encyclopedias (類書). Unlike Western encyclopedias or bibliographies, Leishu organize knowledge through thematic categorization and intertextual quotation, blending literary, historical, and philosophical materials in hierarchical structures. They serve as repositories of knowledge and as mechanisms for the preservation of lost works. The Leishu Ontology proposed in this study fills this gap by modeling compilation logic, content hierarchies, and intertextual relationships as semantic entities. Therefore, this work situates the traditional practice of reconstructing lost texts, within the contemporary frameworks of Digital Reconstructive Compilation, bridging computational representation and philological scholarship.
3.Methodology
This study introduces the Leishu Ontology as a formal semantic model that restructures the knowledge system of ancient Chinese encyclopedias. Following the Stanford seven-step ontology modeling methodology, a top-down approach was adopted to define the conceptual framework and data relations. The ontology comprises 101 classes, 93 object properties, and 44 data properties, encompassing entities such as encyclopedia, compilation style, content structure, and topic, as shown in Figure 1. These semantic entities establish a unified framework that allows multiple encyclopedias to be aggregated under a shared conceptual schema.
Figure 1 the core classes of Leishu Ontology
The research draws on three representative encyclopedias: the comprehensive Yongle Dadian, the literary Pian Zhi, and the specialized Quan Fang Bei Zu. These works span different dynasties, compilation methods, and thematic focuses, providing a balanced dataset for ontology construction. Texts were processed through OCR recognition, manual proofreading, and metadata enrichment to produce structured entries for semantic annotation. Each Leishu’s content structure is decomposed into hierarchical levels (H1–H4) corresponding to volumes, sections, and entries, with quotation books, compilers, and topics modeled as entities to reveal cross-textual relationships. This hierarchical model transforms traditional classification systems into a machine-interpretable semantic network supporting retrieval and inference.
Building on the ontology, we developed a Knowledge Service Platform to operationalize the reconstructed framework. The platform integrates functions including information filtering, image-text comparison, variant text retrieval, semantic search, and reconstruction assistance, with the ontology provideing the core structural framework linking entities and relationships. Semantic retrieval uses sentence-level vectorization and similarity to identify cross-encyclopedia related fragments, while variant retrieval enables precise cross-version textual comparison, supporting philological verification and cross-encyclopedic analysis with high accuracy.
Figure 2 Workflow of the platform
Implementation follows a three-phase workflow (Figure2). By integrating heterogeneous data through ontology mapping, the platform transcends conventional digital database limitations, enabling users to aggregate knowledge units scattered across Leishu. Embedded visualization tools and semantic reasoning transforms it from a static repository into a dynamic research environment supporting textual reconstruction and knowledge discovery.
4.The Leishu Knowledge Service Platform
Building upon the Leishu Ontology, this study develops the Leishu Knowledge Service Platform(http://book.dhontology.com/), which supports the intelligent exploration, retrieval, and reconstructive analysis of ancient Chinese encyclopedias. it semantically integrates heterogeneous Leishu resources within a unified ontology framework, providing modules for variant text retrieval, semantic search, image–text comparison, and textual reconstruction ( Figure 3).
Figure 3 Homepage of the platform
4.1 Variant Text Retrieval
The variant text retrieval module enables scholars to identify and compare textual versions of the same entry across Leishu. Using an Elasticsearch engine with Chinese word segmentation and cosine similarity, it automatically detects substitutions, omissions, paraphrases and additions, supporting textual collation, variant verification, and intertextual comparison. It replaces time-intensive manual collation with a scalable computational process, enhancing precision and efficiency.
For example, querying the Tang poem line “愁來一日知長”(Sorrow comes, and a single day feels long) returns variants like “愁來一日即知長”(Sorow comes, and a single day immediately feels long) and “愁來一日即为長”(Sorrow comes, and a single day thus feels long), ranked by similarity score ( Figure 4), demonstrating the platform’s capability to recognize subtle variations and surface semantically related expressions across encyclopedias for precise comparative analysis.
Figure 4 Case for Variant Text Retrieval
4.2 Textual Reconstruction and Lost Work Recovery
A core innovation of the platform is its textual reconstruction module(輯佚), which reassemble quotations of lost works preserved across encyclopedias. Using semantic clustering, it automatically gathers related fragments to reconstruct dispersed citations into coherent virtual representations of lost texts, which can be visualized and annotated to trace citation networks, transmission pathways, and intertextual relationships. This transforms encyclopedias from static reference into dynamic knowledge systems revealing classical text structure and transmission.
Many encyclopedias contain rich materials from ancient scientific and technical text collections. For instance, the lost agricultural treatise Fan Shengzhi Shu(《氾勝之書》) , compiled several centuries earlier than Qimin Yaoshu, survives only through scattered quotations preserved in works like Taiping Yulan and Yiwen Leiju. By aggregating and aligning these fragments across Leishu, the platform enables virtual reconstruction of its content and structure, recovering an important piece of agricultural knowledge and demonstrating the power of ontology-driven textual reconstruction (Figure 5).
Figure 5 The Case of Fan Shengzhi Shu
5.Conclusion
This study proposes an ontology-driven approach for the digital reconstruction of ancient Chinese encyclopedic knowledge and the historiography of Textual Reconstruction.
The findings demonstrate that the ontology-based model transforms fragmented, heterogeneous Leishu knowledge into an integrated semantic system. The Leishu Ontology captures both the macro-level classification hierarchies and micro-level text quotations, enabling machine-actionable representations of complex philological relationships. The Leishu Knowledge Service Platform operationalizes these relationships through intelligent retrieval and reconstruction assistance, directly addressing textual criticism and data heterogeneity challenges.
*Acknowledgement: This study is supported by the NSFC project "Research on the Construction of Cross-Language and Multimodal Knowledge Graph of Cultural Relics" (Grant No.72204011).
Bleeker, E., Buitendijk, B., Dekker, R.H., Neyt, V., &Hulle,D.V.(2022).Layers of Variation:a Computational Approach to Collating Texts with Revisions.In Digital Humanities Quarterly.
Cai, M. (2023). Ontology construction of traditional Chinese handmade paper:A CIDOC-CRM approach.Digital Scholarship in the Humanities,38(4),1431–1444.
Crane,G., Babeu,A.,&Bamman,D.(2006).Beyond digital incunabula:Modeling the next generation of digital classics.Journal of Digital Humanities,8(1).
Doerr, M. (2003). The CIDOC Conceptual Reference Module: An Ontological Approach to Semantic Interoperability of Metadata. AI Magazine, 24(3), 75.
Doerr, M., Bekiari, C., & Lê, P. (2008). FRBR OO, A CONCEPTUAL MODEL FOR PERFORMING ARTS.
Freire,N.,Robson,G.,Howard,J.B.,Manguinhas,H.,&Isaac,A.Cultural heritage metadata aggregation using web technologies: IIIF, Sitemaps and Schema.org.Int J Digit Libr 21,19–30(2020).
Isaac, A., Charles, V., Fernie, K., Dallas, C., Gavrilis, D., & Angelis, S. (2013). Achieving Interoperability between the CARARE Schema for Monuments and Sites and the Europeana Data Model. Proceedings of the International Conference on Dublin Core and Metadata Applications, 2013.
Mahadevan,A., Mathioudakis,M.,Mäkelä,E.et al.Text reuse in large historical corpora:insights from the optimization of a data science system.Int J Data Sci Anal 20,4631–4643(2025).
Melo, D. 2023. A strategy for archives metadata representation on CIDOC-CRM and knowledge discovery. Semantic Web, 14(3), 553–576.
Pierazzo, E. (2015). Digital scholarly editing:Theories,models and methods.Routledge.
Porter, D., Campagnolo, A., &Connelly,E.(2017).VisColl:A New Collation Tool for Manuscript Studies.In Codicology and Palaeography in the Digital Age 4,81-100,2017.
PROJECT: Fragmentarium:Laboratory for Medieval Manuscript Fragments https://fragmentarium.ms/
PROJECT: Vatican Library https://digi.vatlib.it/
Wang, L., Wei, T., & Wang, J. (2023). Modeling Chinese ancient book catalogs with CABC ontology.In Z.Jin et al.(Eds.),Knowledge Science,Engineering and Management(KSEM 2023)(pp.367–379).Springer.
Wang, L., Wang, J., & Wei,T.(2024).Using ontology to model time description in historical Chinese texts.Digital Scholarship in the Humanities.Advance online publication.
Zheng,X., Li, M., Wan, Z., & Zhang, Y. (2024). Knowledge mining and graph visualization of ancient Chinese scientific and technological documents.Library Hi Tech,42(6),1693–1721.