Daejeon, July 27–31
This curatorial project focuses on Korea’s World Literature Collections (Segye Munhak Jeonjip 세계문학전집) to address a fundamental challenge in bibliographic modeling. The complex, multi-work publishing formats of the World Literature Collections make scholarly use of this culturally rich data extremely difficult, and our goal was to structure the full bibliographical data, along with its metadata, as a relational framework. The corpus we have collected encompasses over 4,500 published instances from more than 145 publishers held at the National Library of Korea, spanning six decades of translation and publication of what has been considered the canon of the world. While the cultural significance of this corpus as a record of Korea’s engagement with global literary traditions is indisputable, its complete data has never been systematically curated, let alone fully available, for scholarly use. Recent scholarship in the relevant field has highlighted the role of translation in Korean canon formation (Cho 2016) and the infrastructures of Korean literary production (Levitt & Saferstein 2022), but these studies rely on qualitative analysis.
Meanwhile, linked data and semantic web technologies have become increasingly important in creating, publishing, and analyzing cultural heritage data in Digital Humanities (Bikakis et al. 2021). And BIBFRAME implementations at major national libraries, including the Swedish National Library’s full transition in 2018 and continued efforts at the Library of Congress, demonstrates the potential of linked data for bibliographic description. However, these implementations have largely focused on Western collections with relatively consistent metadata. The core difficulty of retrieving this data systematically lies in inconsistencies within NLK records, including the routine omission of series statements (which is supposed to mark the works published within a single collection), conflation of distinct editions, and errors accumulated through decades of evolving transliteration standards for foreign names and titles—partly due to the lingering tradition of Korean-Chinese mixed script.
Our approach involves twofold efforts of scope expansion and relational curation. We began by retrieving and cleaning the complete corpus of earliest World Literature Collections records from the 1950s and 60s and manually cross-checked the NLK catalogue data against physical copies of first editions to ensure data consistency and reliability. We then utilized both the computational harvesting of bibliographic data and manual cataloguing of detailed information missing from the catalogue—whether due to inconsistent data management or the structural complexity of anthologies containing multiple works by multiple authors. This multi-step process culminated in adding sub-work listings, ISNI identifiers for authors and translators, and external links to original works. ISNI, as an ISO-certified global standard designed to act as a bridge identifier across multiple domains, is a critical component in Linked Data applications (Guerrini & Possemato 2013); our use of ISNI ensures accurate resolution of the often non-trivial relationships between instances and works—a key requirement for linked data modeling as articulated in the BIBFRAME Work-Instance framework (Library of Congress).
The dataset that we created is, in itself, an ambitious bibliographic project, but it goes beyond historical preservation of cultural material. The dataset opens up new directions for research into Korean cultural history by actively redefining the study of translated literature. Translation studies are no longer limited to isolated textual critiques. With well-organized resources, they can now incorporate rigorous data-driven analyses of cultural transmission and reception. Scholars can quantitatively track how the imported canon formed and shifted over six decades, examining the rise and fall of different national literature’s prominence throughout the tumultuous history of modern Korea. The structured data also enables comparative analysis of publisher curation strategies, revealing how major houses positioned themselves through their selection and presentation of world literature, which in turn shaped Korean readers’ perceptions of world literature. Researchers can also map networks of translator influence, identifying important figures who shaped Korean readers’ access to particular literary traditions. Beyond these immediate applications, the dataset can also provide a testbed for BIBFRAME implementation, offering a replicable model for cultural institutions seeking to transform complex national serials into dynamic, linked data resources for global digital humanities scholarship.
Our poster will include visualizations of major quantitative analysis results that this curated dataset yields, in addition to a detailed description of the whole process of assembling the complete and accessible dataset, which allows us to reimagine the future of libraries and the future of data-oriented cultural heritage and literary studies.