DH 2026

Daejeon, July 27–31

Thu, July 3013:40–15:10S034101-102
Long Paper

Small Data, Deep Meaning: A Multilingual Linked-Data Infrastructure for Classical Sino-Korean Poetry Talks

Christina Han
Wifrid Laurier University, Canada · chan@wlu.ca
Lyndsey Twining
Independent Researcher · lyndseytwining@gmail.com
Jing Hu
Berline State Library · hu777jing@gmail.com
Yeong Won Chi
Korea University · chi123123@naver.com
Jonghoon Yoon
Arts Council Korea Artsarchive · hoonjong0821@gmail.com

This paper presents a multilingual Linked Open Data (LOD) platform based on Semantic MediaWiki (SMW) for the study of classical Sinitic literature, centered on the seventeenth-century Korean anthology Sihwa ch’ongnim 詩話叢林 (Compendium of Poetry Talks). The project introduces a new model for digital curation in the humanities: a small-scale, fully curated, interoperable, multilingual LOD environment that supports interpretive, historically grounded research on classical East Asian texts. It addresses three enduring challenges: the lack of a semantic, multilingual infrastructure for Korean sihwa (poetry talks), the technical inaccessibility of existing LOD tools for non-specialists, and the absence of a replicable platform for managing cross-script humanities data.

Sihwa is a hybrid genre unique to East Asia, blending anecdotal prose with embedded poems and critical commentary. The Sihwa ch’ongnim is the largest surviving Korean sihwa collection and reflects many dimensions of premodern Korean literary culture. Yet its fragmentary structure, non-linear narratives, and dense person–place–concept references make it difficult to analyze through print editions or monolithic databases [Han et al., 2022]. To date, no project has offered a fully semantic, multilingual, cross-script infrastructure for sihwa.

Situating the Project

Recent digital humanities research employs diverse computational methods for analyzing classical Sinitic poetry, from text mining and BERT models to graph analyses in Neo4j [Cai et al. 2024; Ha et al. 2025; Yu et al. 2024; Zhang, Q. et al. 2025; Zhang, W. et al. 2025; Zou et al. 2025]. While these approaches uncover large-scale patterns, they rely on custom pipelines and continue to face challenges posed by linguistic ambiguity and complex structures. New Large Language Model (LLM)- and Retrieval-Augmented Generation (RAG)-based pipelines have expanded possibilities for question-answering and spatio-temporal analysis [Cao, Peng et al. 2024; Cao, Shi et al. 2024; Chen A. et al. 2024; Chen A. et al. 2025; Liu et al. 2024; Ma et al. 2025; Yao et al. 2025; Yoo 2024], yet these tools remain largely monolingual, engineering-focused, and difficult for many humanists to adopt.

Within the LOD ecosystem, SMW has been used only sparingly in literary or cultural projects. The Living Poets project applies SMW to map the reception of Greek and Latin authors, while others—such as the Digital Victorian Periodical Project, the Encyclopedia of Romantic Nationalism, the Jiam Diary, and Mapping the Republic of Letters—depend on custom interfaces, TEI encoding, or proprietary visualization layers. Tools like MARKUS, COMARKUS, and IMMARKUS provide sophisticated environments for annotating East Asian texts and images [De Weerdt et al. 2025], but fully using these systems typically requires expertise in computer languages and specialized infrastructure [Chi & Choi 2024; Doh 2025; Li, S. et al. 2025].

More broadly, LOD tools continue to suffer from significant usability barriers, especially for humanities researchers unfamiliar with graph modeling [Barbera 2013; Zheng et al. 2025]. Evaluations of linked-data adoption highlight the need for accessible interfaces, sustainability, and clear documentation [Middle 2022; Linked Art LOUD].

Our Project

Our platform is designed in direct response to these gaps. Built on SMW, it demonstrates that rich semantic data for classical Sinitic texts can be created, navigated, and reused through an accessible, multilingual interface. It draws on conceptual data modeling for sihwa and classical Korean Sinitic poetry, extends modeling approaches from civil service examination data, and incorporates insights from ontology construction in poetry and narrative studies [Lee, G. et al. 2024; Lee, B. 2024; Liu et al. 2018].

Technically, the platform addresses key barriers in classical Korean literary studies: it assigns stable URIs to sihwa entities previously known only through print citations, structures poems–criticism–people–places as queryable RDF triples rather than isolated text, and links Korean literary figures to databases like Wikidata—enabling cross-referenced biographical context absent from print editions. The implementation follows Tim Berners-Lee’s rules for Linked Data while improving usability in line with Middle’s Five-Star Model and the LOUD (Linked Open Usable Data) principles [Middle 2021; Linked Art LOUD].

Ontology and Multilingual/Cross-Script Design

The project’s ontology is designed to represent the literary and narrative structure of sihwa. Entries, poems, and critiques are modeled as discrete text entities within a hierarchical bibliographic framework. Surrounding these core entities are 9 classes—work, person, place, era, critical term, and topic—supported by 25 datatype properties and 16 object properties. This curated property set ensures conceptual clarity and efficient querying.

The dataset currently includes 8,100+ interconnected entities: 932 entries, 1,811 poems, and 1,743 critiques, linked to 120 literary works, 1,222 people, 547 places, 44 eras, 689 critical terms, and 1,055 topics. This constitutes a fully annotated semantic representation of the entire Sihwa ch’ongnim. Contextual classes continue to expand as new thematic patterns and motifs are identified, illustrating how small-scale curation supports iterative enrichment unavailable in static big-data models.

The ontology captures both bibliographic and interpretive context. Datatype properties manage multilingual names, dates, romanization, and external IDs; object properties encode roles and relationships such as “subject” and “creator.” Properties are grouped—for example, all multilingual name properties are grouped under a super-property—allowing both detailed display and straightforward querying.

Image 1) Entry page showing Sinitic text, related information, network graph, map, and timeline

The platform supports both non-specialists and scholars through multilingual interfaces in English, Korean, Chinese, and French. Each entity provides names in Hanja, Hangul, English, McCune–Reischauer, Revised Romanization, and Pinyin, with texts shown in Classical Sinitic alongside parallel Korean and English translations. This enables multilingual reading, translation comparison, and philological analysis across varying language proficiencies. Crucially, users can search Korean terms while browsing the English interface (and vice versa), lowering barriers between Korean, Sinological, and Anglophone research communities and reflecting real-world multilingual scholarly practice.

Platform Implementation

The platform is built with MediaWiki, Semantic MediaWiki, PageForms, and extensions such as ModernTimeline and KnowledgeGraph. This stack enables semantic data creation through user-friendly forms: when contributors enter information, the system automatically generates semantic annotations, updates filters and timelines, and populates indexes. Dynamic homepage carousels (e.g., “Women,” “War,” “Diplomatic Encounters,” “Food”) are produced entirely from semantic queries.

This form-based design addresses usability concerns raised by Barbera (2013) and Middle (2021)  regarding the technical barriers to Linked Data adoption in the humanities. Instead of requiring SPARQL or RDF knowledge, contributors work in a familiar environment, with semantic relationships inserted invisibly through structured forms—making “thinking in the graph” accessible to humanities scholars [Lin 2020; Si 2026]. Its reliance on open-source software supports sustainability and replicability: institutions can reuse templates, adapt the ontology, or add new collections.

Image 2) Semantic data input form

Image 3) Filtered data search form

In consideration of diverse user technical ability,  the platform offers both visual dropdown filters for basic searches and SPARQL-like querying for complex searches. For LOD integration, each entity links to external URIs when available, particularly those of Korean research institutions, Wikidata, and national LOD repositories. The platform currently provides outbound links; the next stage will import external data via the External Data extension. Because Wikidata often includes multiple authority IDs and images for Korean figures, automated import will enrich biographical context without duplicating labor. We plan to publish our own URIs to Wikidata, enabling bidirectional integration.

Case Study

Our platform allows users to explore how a given concept appears throughout the Sihwa ch’ongnim anthology across temporal, geographic, and conceptual contexts via searchable tables, timelines, maps, and graph visualizations. Such visualizations are created directly on a wiki page using SMW query language. For example, a timeline of texts mentioning the Han River is queried and displayed through the following code: {{#ask: [[hasSubject.nameEng::L037]] |?hasCreator.yearBirth= |?hasCreator.yearDeath= |format=moderntimeline}}. Meanwhile, a map of places appearing alongside the Han River is generated thusly: {{#ask: [[-hasSubject.hasSubject.nameEng::L037]] |format=leaflet |?gis= }}. Going forward, we plan to create forms to further automate this visualization process.

Image 4) Timeline of texts mentioning the Han River

Image 5) Map of places mentioned alongside the Han River

Contributions to DH Research: Small Data, Deep Meaning

This paper advances three core arguments. First, small, fully curated semantic datasets complement big-data and NLP methods by enabling historically grounded interpretation. While large-scale analyses reveal macro-patterns in premodern poetry [Hou & Frank 2015; Yu et al. 2024; Liu et al. 2025; Lee B. 2024; Zhang et al. 2023], interpretive questions require narratively encoded detail [Lee S. 2024; Do 2025].

Second, the project introduces an accessible, multilingual LOD model that targets long-standing usability gaps in DH. It lowers barriers that have limited LOD adoption and reframes linked data as collaborative scholarship rather than technical work [Barbera 2013; Middle 2021].

Third, the platform offers a replicable, sustainable model for classical text projects across East Asia. Built entirely on open-source tools, it can scale to additional sihwa collections or adapt to other genres.

Overall, the project demonstrates that a sustainable, multilingual, cross-script linked-data platform for Sinitic literature can be built using accessible, open tools. By prioritizing interpretive richness and usability, it outlines a blueprint for DH work that moves beyond pattern detection toward meaning-making grounded in structured, shareable data.

References
  1. Barbera, M. (2013): “Linked (open) data at web scale : research, social and engineering challenges in the digital humanities”, in: JLIS.It : Italian Journal of Library and Information Science, 4, 1: 91–104. DOI: 10.4403/jlis.it-6333.
  2. Cai, Z., / Kurzynski, M. (2024): “Brevity and Breadth: A Linguistic, Aesthetic, and DH-Assisted Study of the Book of Poetry and “Nineteen Old Poems”, in: Journal of Chinese Literature and Culture, 11, 2: 235–264. DOI: 10.1215/23290048-11410394. 
  3. Cao, J. / Shi, Y. / Peng, D. / Liu, Y. / Jin, L. (2024): “C 3 Bench: A Comprehensive Classical Chinese Understanding Benchmark for Large Language Models”, in Computation and Language, May 27. DOI:  10.48550/arxiv.2405.17732. 
  4. Cao, J. / Peng, D. / Zhang, P. / Shi, Y. / Liu, Y. / Ding, K. / Jin, L. (2024): “TongGu: Mastering Classical Chinese Understanding with Knowledge-Grounded Large Language Models” in: Computation and Language, September 30. DOI: 10.48550/arxiv.2407.03937. 
  5. Chen, A. / Lou, L. / Chen, K. / Bai, X. / Xiang, Y. / Yang, M. / Zhao, T. / Zhang, M. (2024): “Large language models for classical Chinese poetry translation: Benchmarking, evaluating, and improving” in: Computation and Language, December 30. DOI: 10.48550/arXiv.2408.09945. 
  6. Chen, A. / Lou, L. / Chen, K. / Bai, X. / Xiang, Y. / Yang, M. / Zhao, T. / Zhang, M. (2025): “Benchmarking LLMs for translating classical Chinese poetry: Evaluating adequacy, fluency, and elegance”: in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 33008–33025). Association for Computational Linguistics. DOI:  10.18653/v1/2025.emnlp-main.1678.
  7. Chen, J. / Shi, H. (2024): “Construction of LLM–RAG pipeline based on spatial narrative characteristics of Yanxinglu” in: Segae Hanja Yŏn’gu, 7, 2. 
  8. Chi, Y. W. / Choi, J. K. (2024). “A conceptual data modeling attempt for building a Korean classical poetry database” in: Minjok Munhwasa Yŏn’gu, 85, 43–82. 
  9. De Weerdt, H./ Ho, H. I. / Simon, R. / Lee S. / Molenaar, S. / Xi, W. / Zhuang, D. / Stojević, I. / Tu, H.-C. / Zaneri, T. / Lin, N.-Y. / Meister, M. (2025): “Contextual semantic text and image annotation in the MARKUS environment” in: Digital Humanities Quarterly, 19, 4. <https://dhq.digitalhumanities.org/vol/19/4/000808/000808.html> [15.08.2025].
  10. Digital Victorian Periodical Poetry. <https://dvpp.uvic.ca/index.html> [04.05.2026]
  11. Doh, W. Y. (2025): “The necessity and direction of compiling a dictionary of terms in Korean classical Chinese texts” in: Minjok Munhwa, 69, 117–146. 
  12. Encyclopedia of Romantic Nationalism. <https://ernie.uva.nl/viewer.p/21/56> [04.05.2026]
  13. Gao, Y. / Xiong, Y. / Gao, X. / Jia, K. / Pan, J. / Bi, Y. / Dai, Y. / Sun, J. / Wang, M. / Wang, H. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey” in: Computer and Language (March 27). DOI: 10.48550/arxiv.2312.10997. 
  14. Ha, D. J. / Park, M. (2025). “Text mining analysis of jingpai and haipai literary works: Possibilities and limitations of RAG-based chatbot analysis” in: Simininmunhak, 49, 203–231. DOI: 10.22842/kgucfh.2025.49.203.
  15. Han, Christina. / Chi, Y. W. / Hu, J. / Ryu, I. T. (2022). “A foundational design for creating a sihwa semantic data archive” in: Hanmunhak Nonjip, 63: 105–146. DOI: 10.17260/jklc.2022.63..105. 
  16. Han, H. / Wang, Y. / Shomer, H. / Guo, K. / Ding, J. / Lei, Y. / Halappanavar, M. / Rossi, R. A. / Mukherjee, S. / Tang, X. / He, Q. / Hua, Z. / Long, B. / Zhao, T. / Shah, N. / Javari, A. / Xia, Y. / Tang, J. (2024): “Retrieval-Augmented Generation with Graphs (GraphRAG)”: in Computer and Language (January 8). DOI: 10.48550/arxiv.2501.00309.
  17. Horvath, A. (2021): “Enhancing language inclusivity in Digital Humanities: Towards sensitivity and multilingualism”, in: Modern Languages Open 1. DOI: 10.3828/mlo.v0i0.382.
  18. Hou, Y. / Frank, A. (2015): “Analyzing sentiment in classical Chinese poetry”, in: Association for Computational Linguistics (ed.): Proceedings of the 9th SIGHUM Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities (LaTeCH), July 2015: 15–24. DOI: 10.18653/v1/W15-3703.
  19. Jeong, C. (2024): “A graph-agent–based approach to enhancing knowledge-based QA with advanced RAG”, in: Chisik Kyŏngyŏng Yŏn'gu 25, 3: 99–119https://www.kci.go.kr/kciportal/landing/article.kci?arti_id=ART003120450
  20. Kang, C. (2025): “ChatGPT utilization in the study and education of Geumo Sinhwa: With critical reflections”, in: Ŏmun Ronch'ong 103: 63–90.
  21. Lee, B. C. (2024): “Application and limitations of text mining techniques in classical Sinitic poetry”, in: Ŏmun Yŏn'gu 121: 219–249.
  22. Lee, G. H. / Byun, E. M. / Ryu, I. T. (2024): “Semantic data processing of civil service examination materials in the Chosŏn Dynasty”, in: Han'gukhak Munhwa Yŏn'gu 92: 65–104.
  23. Lee, S. E. (2024): “The genealogy of stories: Reading success narratives in yadam through digital humanities methodology”, in: Journal of Korean Culture 66: 143–180.
  24. Li, S. / Hu, R. / Wang, L. (2025): “Efficiently Building a Domain-Specific Large Language Model from Scratch: A Case Study of a Classical Chinese Large Language Model”. DOI: 10.48550/arxiv.2505.11810.
  25. Li, Z. / Wang, Z. / Wang, W. / Hung, K. / Xie, H. / Wang, F. L. (2025): “Retrieval-augmented generation for educational application: A systematic survey”, in: Computers and Education. Artificial Intelligence 8, Article 100417. DOI: 10.1016/j.caeai.2025.100417.
  26. Liang, L. / Bo, Z. / Gui, Z. / Zhu, Z. / Zhong, L. / Zhao, P. / Sun, M. / Zhang, Z. / Zhou, J. / Chen, W. / Zhang, W. / Chen, H. (2025): “KAG: Boosting LLMs in professional domains via Knowledge Augmented Generation”, in: Companion Proceedings of the ACM on Web Conference 2025: 334–343. DOI: 10.1145/3701716.3715240.
  27. Linked Art (n.d.): “LOUD: Linked Open Usable Data”https://linked.art/loud/ [03.05.2026].
  28. Liu, C. / Wang, D. / Zhao, Z. / Hu, D. / Wu, M. / Lin, L. / Liu, J. / Zhang, H. / Shen, S. / Li, B. / Zhao, L. (2024): “SikuGPT: A Generative Pre-trained Model for Intelligent Information Processing of Ancient Texts from the Perspective of Digital Humanities”, in: Journal on Computing and Cultural Heritage 17, 4: Article 53. DOI: 10.1145/3676969.
  29. Liu, C. L. / Mazanec, T. J. / Tharsen, J. R. (2018): “Exploring Chinese poetry with digital assistance: Examples from linguistic, literary, and historical viewpoints”, in: Journal of Chinese Literature and Culture 5, 2: 276–321. DOI: 10.1215/23290048-7257002.
  30. Lin, S. (2020): “The Study of Premodern Chinese Literature in the Digital Era: New Methods of Quantitative Statistics, Databases, and Visualization Analyses”, in: Library Trends 69, 1: 269–288. DOI: 10.1353/lib.2020.0032.
  31. Liu, Z. / Wan, G. / Zuo, X. / Liu, Y. (2025): “Sentiment analysis of Chinese ancient poetry based on multidimensional knowledge attention”, in: Digital Scholarship in the Humanities 40, 1: 214–226. DOI: 10.1093/llc/fqae069.
  32. Living Poets (n.d.): Living Poets <https://livingpoets.dur.ac.uk/w/index.php/Welcome_to_Living_Poets> [03.05.2026].
  33. Lv, W. / Cao, Q. / Liu, X. (2025): “A multi agent classical Chinese translation method based on large language models”, in: Scientific Reports 15, 1: Article 40160. DOI: 10.1038/s41598-025-23904-0.
  34. Ma, B. / Yao, Y. / Haensch, A.-C. (2025): “Capabilities and evaluation biases of large language models in classical Chinese poetry generation: A case study on Tang poetry”. DOI: 10.48550/arXiv.2510.15313.
  35. Ma, Z. / He, J. / Liu, S. (2020): “Representation of the spatio-temporal narrative of The Tale of Li Wa”, in: PloS One 15, 4: e0231529. DOI: 10.1371/journal.pone.0231529.
  36. Mapping the Republic of Letters (n.d.): Mapping the Republic of Letters <http://republicofletters.stanford.edu/> [03.05.2026].
  37. Middle, S. (2021): Investigating Linked Data Usability for Ancient World Research. ProQuest Dissertations & Theses.
  38. Si, Z. (2026): “Mapping the spatial–temporal evolution of imagery in Tang poetry: A computer vision and GIS-based approach”, in: Future Digital Technologies and Artificial Intelligence 2, 1: 1–6https://orcid.org/0009-0008-5936-3849
  39. Spence, P. J. / Brandao, R. (2021): “Towards language sensitivity and diversity in the Digital Humanities”, in: Digital Studies 11, 1. DOI: 10.16995/dscn.8098.
  40. Yao, X. / Wang, M. / Chen, B. / Zhao, X. (2025): “WenyanGPT: A Large Language Model for Classical Chinese Tasks”. DOI: 10.48550/arxiv.2504.20609.
  41. Yoo, J. J. (2024): “A preliminary study on the construction of an ontology for the study of classic Sinitic poetry in late Chosŏn Korea: Focusing on the function of inference”, in: Hangukhak Nonjip 96: 5–41.
  42. Yu, H. / Gao, C. / Li, X. / Zhang, L. (2024): “Ancient Chinese Poetry Collation Based on BERT”, in: Procedia Computer Science 242: 1171–1178. DOI: 10.1016/j.procs.2024.08.179.
  43. Zhang, Q. / Chen, S. / Bei, Y. / Yuan, Z. / Zhou, H. / Hong, Z. / Chen, H. / Xiao, Y. / Zhou, C. / Dong, J. / Chang, Y. / Huang, X. (2025): “A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models”. DOI: 10.48550/arxiv.2501.13958.
  44. Zhang, W. / Lai, Z. / Tang, S. (2025): “Spatiotemporal distribution characteristics of Nanjing place names—Based on data mining of Tang-Song poetry and online travelogues”, in: PloS One 20, 2: e0319244. DOI: 10.1371/journal.pone.0319244.
  45. Zhang, W. / Wang, H. / Song, M. / Deng, S. (2023): “A method of constructing a fine-grained sentiment lexicon for the humanities computing of classical Chinese poetry”, in: Neural Computing & Applications 35, 3: 2325–2346. DOI: 10.1007/s00521-022-07690-8.
  46. Zheng, M. / Moeller, S. (2025): “Challenges in processing Chinese texts across genres and eras”, in: Association for Computational Linguistics (ed.): Proceedings of the 9th Widening NLP Workshop, November 2025: 230–234.
  47. Zou, D. / Chen, Y. / Li, M. / Miao, S. / Liu, C. / Han, B. / Cheng, J. / Li, P. (2025): “Weak-to-Strong GraphRAG: Aligning Weak Retrievers with Large Language Models for Graph-based Retrieval Augmented Generation”. DOI: 10.48550/arxiv.2506.22518.