DH 2026

Daejeon, July 27–31

Thu, July 3016:30–18:00S078101-102
Short Paper

From Scattered Sources to Semantic Networks: A Data Curation Methodology for Korean Modern-Era Materials

Jisun Kim
Research Institute of History and Culture, Duksung Women's University, Korea, Republic of (South Korea) · jisundh@gmail.com

Korea's modern-era materials present unique challenges for researchers: they are vast in scale, diverse in type, and scattered across multiple digital archives that function primarily as image repositories.

Ryu (2024) points out the limitations of existing digital archives, which remained focused on ‘digital conversion and provision.’ They failed to move from passive conservation (mere existence) to proactive preservation (active utilization). This suggests the transition to a data archive, pursuing the ‘construction and provision of interactive and systematic data,’ is not yet achieved.

While these archives improve access to original sources, they leave researchers with the burden of independently gathering materials, extracting information, and establishing connections. This research addresses this challenge by presenting a concrete implementation of a semantic data archive that structures scattered sources into a unified knowledge network through meticulous data curation and modeling, demonstrating how individual researchers can conduct systematic data-based humanities research.

Research exploring the transition from repositories providing mere access to sources to systems implementing knowledge sharing archives has been consistently conducted by the Cultural Informatics program at AKS. Relevant studies include dissertations such as (Kim, Ba-ro 2017; Kim, Ji-myeong 2017; Jeong 2018; Kim 2019; Lee 2021).

Case Study: Whashin Department Store (1932-1937)

This research examines Whashin (和信) Department Store in Jongno, during its initial operational period (1932–1937). Among the five major department stores of 1930s Gyeongseong—four of which were Japanese-owned (Hirata[平田], Chojiya[丁子屋], Mitsukoshi[三越], and Minakai[三中井])—Whashin stood as the only Korean-owned establishment. Located in Jongno, the center of Korean residential and commercial life, it functioned not merely as a retail space but as a symbol of Korean commercial autonomy and cultural identity in colonial Korea. Whashin thus serves as an ideal case study for exploring diverse source types within focused scope. The research compiled newspaper articles, advertisements, photographs, diaries, and literary works, examining how factual documentation and fictional representation intersect in cultural memory.

Data Collection and Curation Process

The data collection process began with keyword searches across multiple digital archives. For newspaper articles, 2,186 articles were finally selected from an initial pool of 4,421 results following a meticulous review. This disambiguation process—distinguishing the Jongno flagship store from homonyms (e.g., 和新商店, 畵伸) and branch store references—ensured the spatial specificity of the dataset through manual verification. Each article was catalogued with detailed metadata (publication date, page number, section, headline). For literary works, the research analyzed 31 novels published between 1932–1937 that featured Whashin Department Store, examining 123 specific scenes where the department store appeared as a setting, plot element, or cultural reference.

Ontology Design and Data Modeling

The research designed a domain ontology comprising 25 classes organized into seven core categories: Sourceful Elements (WebResource, Archive, Institution), Notional Elements (Novel, Article, Diary, Newspaper, Magazine, Press), Textual Elements (Photo, Text, Entry, Drawing), Factual Elements (DepartmentStore, Person, Group, Place, Object, Term), Constructed Elements (Event), Narrative Elements (Scene, Character, FictionalSpace, Act), and Relational Elements (Time). This structure systematically models both documentary sources and their informational content while maintaining crucial distinction between factual documentation and fictional representation.

The ontology draws on three frameworks: EKC data model

The EKC Data Model was initially designed by the Digital Humanities Research Institute at AKS in 2016–2017 for Development of Digital Storytelling Resources. It has since been continuously expanded through research on semantic data archives.

for encyclopedic archive implementation, EDM

Europeana Data Model is a data model developed for Europeana, the European Union's digital cultural heritage platform launched in November 2008. Europeana aggregates digitised cultural heritage resources from museums, libraries, archives, galleries, and universities across Europe. EDM was designed to transcend domain-specific metadata standards while preserving each institution's unique metadata structure, enabling semantic interoperability across distributed heritage institutions.

for distributed resource handling, and CIDOC CRM

CIDOC Conceptual Reference Model is a conceptual model designed for the semantic integration of cultural heritage information scattered across museums, libraries, archives, and other institutions.

for event-centric modeling. A key design decision treated DepartmentStore as both agent and place, reflecting its dual nature as commercial actor and physical space. Relationships were precisely defined: DepartmentStore, Group, and Person entities connect to Event entities through specific roles, while Event entities connect to spatial and temporal contexts. The resulting dataset comprises approximately 8,000 records, enabling complex queries tracing how documented events were simultaneously reported in newspapers and fictionalized in novels.

IIIF Image Processing

The International Image Interoperability Framework (IIIF)

IIIF is an international standard framework developed from 2011, led by the British Library and Stanford University, enabling institutions to share and interoperate digitised resources across distributed archives through a standardised set of APIs. The IIIF Consortium was formally established in 2015.

enabled multi-layered image annotation. A 1933 photograph of Kim Ok-jeong at Whashin's toy department could be annotated at multiple levels: the entire photograph, then nested annotations identifying person and location as distinct regions. While IIIF provides technical standards for structuring image data, the ontology defines semantic relationships—that Kim was employed there, that the toy department belongs to Whashin, and how these entities connect across sources. This combination enables systematic connection of visual and textual information across distributed archives.

Fact-Fiction Spectrum Implementation

The fact-fiction spectrum addresses how to connect documentary sources with literary representations within the unified data structure. Each source was categorized by its position on the spectrum: newspaper articles as factual documentation, and novels as fictional representation. Specific events, spaces, and objects could appear across this spectrum. The data structure captures how events documented in newspaper articles were simultaneously fictionalized in novels depicting those events, enabling queries like “show me real events and the novels that depicted them.”

Neo4j Implementation and Query Results

The dataset was implemented in Neo4j graph database, comprising approximately 8,000 records across 25 classes. Factual data queries demonstrated analytical power: context analysis identified 2,034 event relationships with six distinct roles and 1,046 product relationships across four types. Literary data queries uncovered representation patterns: behavioral analysis identified 15 activity types, with encounter (16) as dominant, followed by browsing (12) and purchasing (11)—suggesting the department store functioned primarily as social space in fiction. Spatial analysis showed Whashin Department Store appeared 71 times across 24 novels, while in front of Whashin appeared 14 times across 12 novels.

Most significantly, cross-domain queries traced connections between documented events and fictional representations. A December 1934 newspaper report of childbirth in a streetcar near Whashin appeared fictionalized in a novel published 14 days later, demonstrating how the fact-fiction spectrum enables discovery of real-world events inspiring literary works.

Engagement and Implications

All data has been processed into CSV and JSON formats for open sharing. The semantic data archive, encompassing the curated semantic datasets and data visualization outputs, is publicly accessible via the author's website (https://jisunkim.info/). Though executed by one researcher, standardized formats enable this work to function as a component of larger knowledge networks. The archive is designed as an extensible foundation, inviting expansion through community participation. This research demonstrates that individual scholars can actively shape digital research infrastructures without institutional support, offering a concrete pathway for humanities researchers to build collaborative knowledge systems in an era increasingly centered on AI and data-driven methods.

References
  1. Academy of Korean Studies, Digital Humanities Research Institute (2017): “EKC Data Model-Draft 1.1.” Korean Documentary Heritage Encyves. https://dh.aks.ac.kr/Encyves/wiki/index.php/EKC_Data_Model-Draft_1.1 [15.12.2025].
  2. Academy of Korean Studies, Digital Humanities Research Institute (2022): “Ontology:EKC 2022.” Hanyangdoseong Time Machine Semantic Data Archive. https://dh.aks.ac.kr/hanyang2/wiki/index.php/Ontology:EKC_2022 [15.12.2025].
  3. CIDOC CRM SIG (n.d.): “What is the CIDOC CRM?” CIDOC CRM. https://cidoc-crm.org/ [15.12.2025].
  4. Cramer, Tom (2011): “The International Image Interoperability Framework (IIIF): Laying the Foundation for Common Services, Integrated Resources and a Marketplace of Tools for Scholars Worldwide.” CNI. https://www.cni.org/topics/information-access-retrieval/international-image-interoperability-framework [15.12.2025].
  5. Cramer, Tom (2015): “IIIF Consortium Formed.” IIIF. https://iiif.io/news/2015/06/17/iiif-consortium/ [15.12.2025].
  6. Europeana Foundation (2017): “Definition of the Europeana Data Model v5.2.8.” https://pro.europeana.eu/files/Europeana_Professional/Share_your_data/Technical_requirements/EDM_Documentation/EDM_Definition_v5.2.8_102017.pdf [15.12.2025].
  7. Jeong, Ju-young (2018): “A Study on the Construction of a Semantic Archive of 1970s Small Theater Plays: Focused on Ejeot-to Warehouse Theater (1975) and Samil-ro Warehouse Theater (1976–1979).” Ph.D. diss., Graduate School of Korean Studies, Academy of Korean Studies.
  8. Kim, Ba-ro (2017): “Archiving and Utilizing Relational Data on Institutional Systems & Personnel Management: Records from Korea's Public Schools in Modern Era (1895-1910).” Ph.D. diss., Graduate School of Korean Studies, Academy of Korean Studies.
  9. Kim, Hyun kyu (2019): “A Study on Constructing Linked Open Data about the March First Movement.” Master's thesis, Graduate School of Korean Studies, Academy of Korean Studies.
  10. Kim, Ji-myeong (2017): “Developing a Digital Archive and Curation Model for Historical Documents: Focusing on the National Debt Repayment Movement in Korea.” Ph.D. diss., Graduate School of Korean Studies, Academy of Korean Studies.
  11. Lee, Sumin (2021): “Implementation of the Meta Archive for Digital Reproduction of the 1929 Joseon Exposition.” Master's thesis, Graduate School of Korean Studies, Academy of Korean Studies.
  12. Ryu, Intae (2024): “From Conservation of ‘Existence’ to Preservation of ‘Utilization’: On the Transition from Digital Archive to Data Archive”, in: Seoul Museum of History Research Division (ed.): 2024 Seoul History Museum Review: Urban History and Archive. Seoul: Seoul Museum of History 58–75.