Daejeon, July 27–31
The Korean Union Catalog System for Old and Rare Materials (KORCIS), established in 2005, achieved its initial goal of integrating bibliographic information on Korean ancient texts at the national level. Twenty years later, however, its structural limitations have become clearly apparent. KORCIS is essentially a digital transplantation of the paper card catalog: while it faithfully describes individual materials as discrete objects according to KORMARC rules, it fundamentally fails to accommodate the relation-centered knowledge modeling that contemporary digital humanities research demands. An even more pressing issue is that the role of cultural heritage data is being redefined in the age of AI. As structured cultural data becomes an essential resource for training large language models (LLMs) and multimodal AI systems, the structural deficiencies of Korean ancient text metadata act as a barrier that undermines the global visibility of Korean cultural knowledge. This asymmetry reinforces the algorithmic bias whereby non-Western cultural knowledge systems remain structurally underrepresented in AI training data.
This study systematically analyzes the metadata management framework of KORCIS using the internationally recognized FAIR principles (Findable, Accessible, Interoperable, Reusable) as a diagnostic framework, and proposes directions for semantic standardization by comparing it with similar systems in Japan and China.
The FAIR principles were originally developed for scientific data management (Wilkinson et al., 2016) but have recently been increasingly applied to cultural heritage and archival management. This study collected and comprehensively analyzed the entire KORCIS bibliographic dataset—567,212 records from 142 participating institutions—released by the National Library of Korea through the Public Data Portal. Using the 15 sub-criteria of the FAIR principles proposed by Wilkinson et al. (2016) as evaluation criteria, compliance was classified into four levels: compliant, partially compliant, insufficient, and absent. This assessment was carried out by synthesizing ① the structural characteristics of the data, ② the availability of system-level functional support (e.g., APIs, SPARQL endpoints, and persistent identifiers), and ③ quantitative indicators such as the occurrence rate of each MARC tag and the rate of authority number entry. The comparative analysis was conducted through the official documentation and web services of Europeana, Japan’s National Diet Library (NDL), and the National Library of China’s Database of Chinese Ancient Books (中華古籍資源庫).
The comprehensive analysis found that not a single one of the 15 FAIR criteria received a fully compliant rating: 10 items (66.7%) were rated insufficient or absent, and 5 items (33.3%) partially compliant. While the completeness of mandatory bibliographic fields—such as title and statement of responsibility (245), physical description (300), and publication statement (260)—exceeded 99%, completeness of optional fields was markedly lower, at 27.35% for subject headings (650) and just 1.49% for electronic location and access (856). Institution-by-institution comparison revealed pronounced disparities in metadata quality, stemming from inconsistent practices in which institutions entered the same information into different fields.
Authority control proves even more problematic. Among records containing personal-name fields (100/700), only 0.73% and 0.48%, respectively, include an authority record control number, and not a single URI from an international authority system such as VIAF, Wikidata, or LCNAF appears anywhere across the full set of 567,212 records. This reflects a structure in which AI systems cannot recognize or reconcile the linguistic and semantic complexity distinctive to Korean ancient texts—the many variant forms of personal names (given name, pen name, posthumous title, temple name), the mixed use of Chinese characters, idu, and Hangeul, and multiple designations for the same individual, such as Zhu Xi (朱熹) and Zhuzi (朱子), or Yi Hwang (李滉) and Toegye (退溪)—that AI systems are structurally unable to recognize and integrate. Because the internal control numbers of bibliographic records do not follow internationally recognized persistent-identifier formats such as HTTP URIs, DOIs, or ARKs, direct reference from external systems is impossible (F1 not met), and because no APIs or SPARQL endpoints are provided, no machine-accessible environment has been established at all (A1 insufficient). As a result, KORCIS data remains disconnected from the global Linked Open Data ecosystem, circulating only within domestic systems. The analysis data and scripts are publicly available in the following repository (https://github.com/SHINEUNSUN1027/korcis-fair-analysis).
Europeana demonstrates a practical model of FAIR implementation: it integrates heterogeneous institutional data into a single ontology via the Europeana Data Model (EDM), assigns HTTP URI-based unique identifiers, applies CC0 licensing, and has built REST API and SPARQL endpoints that link to VIAF, DBpedia, and Wikidata via owl:sameAs (Isaac, 2013). Without undertaking a full system rebuild, the NDL has adopted an incremental transformation strategy that prioritizes releasing authority data in RDF format and operates a SPARQL endpoint (https://id.ndl.go.jp/auth/ndla) linked to VIAF URIs via owl:sameAs. This offers a directly applicable reference model for a realistic transformation path under resource-constrained conditions.
The Database of Chinese Ancient Books, by contrast, freely provides more than 33,000 digitized classical texts, yet has not achieved LOD transformation or international authority linkage. It exhibits problems structurally identical to those of KORCIS: a lack of rich, detailed metadata, inadequate data provenance management, and the absence of construction standards. This shows that LOD transformation and international authority linkage remain a shared challenge across East Asian ancient text metadata more broadly.
The comparative analysis quantitatively confirms that the quality of metadata structure determines the AI representation of cultural knowledge. The finding of zero international authority linkages demonstrates that, unless Korea’s historical figures, places, and events exist as structured entities connected by URIs, AI systems cannot learn the relationships among them—a result that leads to the structural marginalization of non-Western cultural knowledge within the global knowledge network (D’Ignazio & Klein, 2020). Accordingly, this study proposes converting KORCIS from KORMARC/XML to an RDF-based semantic model. The core directions are: ① assigning HTTP URI-based persistent identifiers to all bibliographic and authority entities; ② applying SKOS-based controlled vocabularies (alongside linguistic and semantic standardization addressing issues such as personal-name variants and the polysemy of Chinese characters); ③ linking domestic authority identifiers (KAC) to VIAF and Wikidata URIs via owl:sameAs, thereby securing global connectivity while preserving the domestic authority system; ④ providing public SPARQL endpoints and REST APIs; and ⑤ applying open licenses such as CC0.
The implementation roadmap consists of three phases: ① a pilot phase prioritizing the RDF release of authority data and the development of a KORMARC-RDF converter (2026–2027); ② full data migration and the linking of domestic authority data (2028–2029); and ③ the expansion of public services and international linkages (2030–2031). The project is currently focused on developing the Phase 1 prototype. The design of the RDF-based ontology developed for this purpose (KoDEX: Korean Old Document Extended schema) and the KORMARC-RDF conversion guidelines will be addressed separately in follow-up research.
Using empirical data from 567,212 records, this study demonstrates that the problems facing KORCIS stem not from technical obsolescence but from a fundamental paradigmatic mismatch. An RDF transformation grounded in the FAIR principles offers the most reliable path toward integrating Korean ancient texts into the global knowledge network and securing fair representation of Korean cultural memory in the age of AI.
Keywords: FAIR principles, Korean Union Catalog System for Old and Rare Materials (KORCIS), ancient text metadata, Linked Open Data, RDF, semantic transformation, algorithmic bias, East Asian digital humanities