DH 2026

Daejeon, July 27–31

Thu, July 3015:20–16:20S086105
Short Paper

Building a Neo4j Knowledge Graph Platform for Comparative Studies of East Asian Analects Commentaries (1): Zhu Xi's Lixue and Qian Shi's Xinxue

Seoyun Kim
Pusan National University, Korea, Republic of (South Korea) · sy527991@gmail.com

This study represents the first phase of building a Neo4j-based knowledge graph platform for systematically comparing and analyzing philosophical controversies within the East Asian Analects commentary tradition. While existing digital humanities resources for East Asian classical philosophy have focused primarily on bibliographic and textual databases, cases of encoding inter-commentary philosophical divergences as structured relational data remain rare. This study addresses that gap by focusing on visualizing and analyzing the interpretive structures of the two major schools of Song dynasty Neo-Confucianism: Zhu Xi’s (朱熹, 1130–1200) Lixue (理學, School of Principle) and Qian Shi’s (錢時, 1175–1244) Xinxue (心學, School of Mind).

Qian Shi, a second-generation disciple of Lu Jiuyuan (陸九淵, 1139–1192), is the only figure in the Lu school lineage to have left systematic commentaries on the Four Books including the Analects, making him a pivotal figure for empirically analyzing Xinxue’s hermeneutical methodology. By comparing Zhu Xi’s Lunyu Jizhu (論語集注) and Qian Shi’s Lunyu Guanjian (論語管見), it becomes possible to structurally examine how two philosophical systems construct divergent interpretive logics from the same canonical text.

Data Design and Preprocessing

The full texts of Zhu Xi’s Lunyu Jizhu and Qian Shi’s Lunyu Guanjian are first encoded as XML following TEI (Text Encoding Initiative) standards. TEI XML represents the structural characteristics of classical Chinese commentary texts — the layered relationship between canonical text (經文) and commentary, patterns of citation, and the contextual use of philosophical terminology — in a formalized manner, thereby ensuring interoperability with existing digital humanities resources such as the Korean Classics Comprehensive Database and the Korean Classics Studies Data Collection Database. Based on the resulting TEI XML data, an Excel schema is constructed comprising node-type sheets and an integrated Edge sheet for relationships. Excel functions automatically generate Cypher queries within each sheet, which are then imported directly into Neo4j.

Neo4j Knowledge Graph Implementation

The knowledge graph consists of six node types — Passage, Chapter, Scholar, Commentary, Concept, and Theme — and relationships including hierarchical structure (CONTAINS, BELONGS_TO), thematic connection (RELATED_TO), and interpretive act (WRITES, INTERPRETS, EMPHASIZES, DISAGREES_WITH). Each Scholar node connects to a Commentary node via the WRITES relationship, and the Commentary node in turn INTERPRETS a Passage node, thereby clearly separating the interpretive agent from the interpretive act. Cypher queries enable the identification of divergence patterns in interpretation: for instance, passages where Zhu Xi emphasizes li (理, principle) while Qian Shi foregrounds xin (心, mind) can be extracted, and it becomes possible to track whether such differences cluster around specific themes or distribute across the text as a whole. This reveals structural patterns difficult to detect through traditional close reading. The validity of the graph will be evaluated through expert review of query results and comparison with established scholarly interpretations.

Current Status and Future Applications

The study is currently at the stage of methodology design and chapter-level data structure planning. The implementation will proceed in sequence: TEI XML encoding, Excel schema construction, and Neo4j graph population.

The comparative scope will expand in stages. Phase 2 will add Li Zhi (李贄, 1527–1602), Park Se-dong (朴世堂, 1629–1703), and Itō Jinsai (伊藤仁斎, 1627–1705) to compare Analects interpretations across East Asia in the seventeenth century; Phase 3 will incorporate Jeong Yag-yong (丁若鏞, 1762–1836) and Liu Baonan (劉寶楠, 1791–1855) to encompass nineteenth-century Korean and Chinese philological interpretations. The platform will thereby enable systematic comparative analysis of the Analects interpretive landscape across seven East Asian Confucian scholars spanning the twelfth to nineteenth centuries. By presenting a viable methodology for encoding philosophical interpretation as structured, computationally tractable data, this study contributes to the development of digital humanities infrastructure for East Asian philosophy. The resulting knowledge graph can further serve as a structured knowledge base for Retrieval-Augmented Generation (RAG) systems in classical Chinese philosophy, providing the foundational infrastructure for such future applications.

Keywords: knowledge graph, Neo-Confucianism, commentary comparison, Neo4j, TEI XML, Qian Shi (錢時), Analects

References
  1. Byeon, Eunmi / Lee, Donghak (2025): “Semantic Data Processing of Sangseogohun and Sangseogoju, Part 2: Ontology Design and Parsing”, in: Journal of Humanities, Seoul National University 82, 2: 219–268.
  2. Byeon, Eunmi / Lee, Donghak / Ryu, Intae (2024): “Semantic Data Processing of Sangseogohun and Sangseogoju, Part 1: XML Data Design and Compilation”, in: HANMUNHAKRONCHIP: Journal of Korean Literature in Chinese 69: 287–338.
  3. Byeon, Eunmi / Lee, Gilhwan (2025): “Study on the Digitization of the Collection of Examination Articles (科文選集) Using XML”, in: East Asian Journal of Sinology 20: 153–183.
  4. Han, Christina / Won, Chi Yeong / Hu, Jing / Ryu, Intae (2022): “A Foundational Design for Creating a Sihwa (詩話) Semantic Data Archive”, in: HANMUNHAKRONCHIP: Journal of Korean Literature in Chinese 63: 105–146.
  5. Jang, Moon-seok / Ryu, Intae (2021): “Digital Humanities and Study of Korean Literature (1): Semantic Database Design for Author Research”, in: Journal of Korean Literary History 75: 347–426.
  6. Kim, Seoyun (2024): “Analysis of the Korean Confucian Database XML Document and Text Data Design”, in: Daedong Munhwa Yeon’gu 128: 201–236.
  7. Kim, Seoyun (2025): “XML Data Modeling of Analects Commentaries for Semantic Database Construction”, in: HANMUNHAKRONCHIP: Journal of Korean Literature in Chinese 71: 211–242.
  8. Lee, Gilhwan (2024a): “A Preliminary Study on Semantic Data Modeling for Korean Novels in Literary Sinitic: Considerations for Building a Database of Korean Novels in Literary Sinitic”, in: Journal of Korean Literary History 85: 83–121.
  9. Lee, Gilhwan (2024b): “A Study on the Construction of a Semantic Data Archive for Korean Novels in Literary Sinitic: A Concrete Design for Building a Korean Novels in Literary Sinitic Database”, in: Han Mun Hak Bo 51: 67–136.
  10. Lee, Gilhwan / Byeon, Eunmi / Ryu, Intae (2024): “Semantic Data Processing of Civil Service Examination Materials in the Joseon Dynasty”, in: Journal of Korean Literature in Classical Chinese 92: 65–104.
  11. TEI Consortium (eds.) (2023): TEI P5: Guidelines for Electronic Text Encoding and Interchange. Version 4.6.0. <https://www.tei-c.org/release/doc/tei-p5-doc/en/html/> [27.04.2026].
  12. Yoo, Jamie Jungmin (2018): “Methods in Digital Humanities and the New Formalism: Experiments of the Stanford Literary Lab”, in: HANMUNHAKRONCHIP: Journal of Korean Literature in Chinese 49: 77–98.