Daejeon, July 27–31
1. Introduction
Historical research publications have largely been examined one by one, evaluated for their arguments, use of sources, and interpretive validity. This micro-level approach is effective for assessing the scholarly merits of individual works, but it has inherent limitations when it comes to understanding the broader trajectory of decades of accumulated scholarship: how the field has collectively shifted its orientations, which themes and concepts have been repeatedly chosen, and how particular concerns have emerged or been reconfigured across periods. When scholarship accumulates at scale — as it has in Korean historiography, where the primary bibliographic record now exceeds 245,000 items — individual close reading is no longer a sufficient instrument for grasping field-wide structural change.
The central research question of this paper is as follows: can historical research publications be systematically analyzed as indicators of the temporal consciousness of Korean historiography? Temporal consciousness here does not refer to a single ideology or historical worldview shared by scholars at a given moment. Rather, it denotes a structure of collective attention — the patterns by which Korean historiography has repeatedly invoked certain pasts, organized its objects of inquiry through particular concepts, and foregrounded specific themes over time. The aim is not to reinterpret individual works, but to analyze long-term shifts in the intellectual landscape of Korean historiography by treating its accumulated publications as a dataset and reading them at scale.
To this end, the paper converts the Bulletin of Korean Historical Studies (BKHS) into a large-scale bibliographic dataset and applies a distant reading methodology. It further recognizes a core limitation of simple keyword frequency analysis — the problem of absolute frequency inflation driven by corpus growth — and addresses this through an expansion across three levels of analysis: word-level, relational-level, and document-level. By combining these complementary analytical layers, the paper argues that distant reading can make the temporal consciousness of a scholarly community legible in ways that close reading traditions cannot achieve alone.
2. Historiographical Background and Methodological Position
Korean historiography has long sustained institutionalized practices of trend reviewing, most visibly through retrospect-and-prospect essays produced biennially and written by period and field specialists. These reviews provide expert micro-level contextualization of recent scholarship within bounded subfields, and they remain an indispensable resource for understanding research developments. However, their genre logic tends to reproduce temporal and disciplinary compartmentalization. Because they are organized around particular eras, themes, and subfields, they are structurally limited in their capacity to capture field-wide reconfigurations that cut across domains and unfold over decades. This constraint becomes particularly acute at scale: the cumulative volume recorded in the BKHS alone continues to grow by several thousand items per year, making comprehensive micro-level coverage increasingly difficult.
Distant reading in Digital Humanities offers a methodological complement to this tradition. As Moretti (2013) proposed, distant reading does not mean the abandonment of reading but the construction of a new scopic position from which distributions, repetitions, shifts, and clusters across large corpora become observable, enabling analysis at a scale that close reading alone cannot reach. Building on this rationale, Scholz (2021) demonstrated a directly relevant historiographical precedent by applying distant reading to a large historical bibliography and reconstructing long-term intellectual change from metadata alone. This paper adopts that approach as a direct methodological template and applies it to the BKHS.
At the same time, consistent with methodological cautions in text mining for historical research (Guldi 2023), this paper treats macro-patterns as evidential signals rather than deterministic explanations. Distant reading results are interpreted in dialogue with existing close reading-based historiographical reviews, and key patterns are returned to targeted close reading for validation. The relationship between the two modes of analysis is therefore one of complementarity rather than competition.
3. Dataset: Structure and Scale of the Bulletin
The dataset for this study is the Bulletin of Korean Historical Studies (BKHS), compiled by the National Institute of Korean History (NIKH), a government-operated historical compilation institution in the Republic of Korea. Since 1973, the BKHS has been published four times a year, continuously surveying and organizing domestic and international scholarship on Korean history. It is characterized by regular publication, institutional curation, and a standardized record structure, which together support its value as a longitudinal metadata corpus rather than a one-off list.
Using the list-download function and standardized record ID structure, the study iterated all registered IDs and retrieved the full set of accessible records, excluding a small number of IDs with unretrievable content. The analytical scope spans 1973 to 2025. While the initial data collection was conducted in 2024 and covered records through that year, 2025 records have been incorporated into the final analysis. As the 2020s do not yet constitute a complete decade, they are not compared directly with the 1970s–2010s on the basis of absolute frequency; instead, normalized frequencies and relative proportions serve as the primary interpretive basis for this period. To our knowledge, no prior study has attempted to convert the BKHS in its entirety into a dataset for macro-level analysis, making this paper a novel methodological contribution within Korean historiography.
Data curation followed explicit cleaning rules: removal of HTML artifacts and appended service labels (e.g., KCI/RISS), derivation of publication year, normalization of page counts, and standardization of publication place across Hangul, Hanja, and Romanized forms, reflecting the known complexity of multilingual and diachronic place-name variation in the Bulletin metadata. Key fields including title, author, publication year, publication type, venue, keywords, and table of contents are largely complete, though coverage varies across fields.
4. Methodology
The methodological design takes simple keyword frequency analysis as its baseline and extends it through three levels of analysis, each targeting a distinct dimension of the data. This multi-level structure is motivated by the recognition that frequency counts alone conflate genuine shifts in scholarly emphasis with the mechanical effect of overall corpus growth.
The first is word-level analysis. TF-IDF is applied to extract relatively distinctive terms from each decadal corpus, identifying vocabulary that is prominent in a given period but not uniformly distributed across the entire corpus. This corrects for the inflation of absolute counts as the corpus grows and allows the identification of terms that are genuinely characteristic of a given decade rather than simply common overall. N-gram analysis complements this by capturing compound research terms — such as singminji gundaeseong (colonial modernity) or Iljegangjeomsgi (Japanese colonial period) — that cannot be captured through single-word analysis alone and that function as coherent thematic units in Korean historiographical vocabulary.
The second is relational-level analysis. Co-occurrence analysis examines which words appear together with a given term across the corpus, tracing shifts in the semantic context in which individual terms are embedded rather than simply tracking their frequency. This distinction is crucial: the same term can carry very different historiographical significance depending on the conceptual network in which it is placed. For instance, minjok (ethnic nation) carries different implications when co-occurring with terms such as independence, liberation, and movement than when it appears alongside identity, memory, and East Asia. Co-occurrence analysis makes these contextual shifts visible and allows the study to interpret keyword trajectories as evidence of conceptual recontextualization rather than simple rise or decline. Both word-level and relational-level analyses operate on sparse, count-based representations of the corpus.
The third is document-level analysis. Embedding-based topic clustering employs dense semantic embeddings and groups semantically similar records into thematic clusters, discovering latent subject structures inductively from within the data without relying on researcher-selected keywords. This constitutes a methodologically distinct paradigm from the previous two levels: whereas TF-IDF, n-gram, and co-occurrence analysis operate on sparse count-based representations, topic clustering works with dense vector representations that capture semantic similarity across documents. These two paradigms are not mutually substitutable but function as complementary analytical layers. The study combines titles, keywords, and tables of contents into a single analytical text per record, constructs topic clusters from this combined representation, and analyzes how the relative weight of each cluster shifts across annual and decadal bins.
The results of this multi-level analysis are interpreted in dialogue with existing historiographical reviews and trend studies. By distinguishing where data-driven patterns align with, diverge from, or add to existing accounts of Korean historiographical development, the study constructs a macro-level verification layer that connects distant and close reading.
5. Findings
This study takes long-term patterns identified through frequency-based analysis as its baseline and extends them through word-, relational-, and document-level analysis. This allows the study to move beyond simple frequency trends and examine conceptual recontextualization and long-term shifts in thematic structure that raw counts alone cannot reveal.
The four indicator terms — minjung (popular masses), minjok (ethnic nation), women, and environment — do not constitute an exhaustive inventory of Korean historiographical concerns. They function as seed terms: concepts repeatedly identified as central in existing historiographical reviews and trend studies. The study cross-examines these seed terms against TF-IDF outputs, co-occurrence results, and topic clustering to clarify the relationship between researcher-selected terms and data-driven distinctive vocabulary, and to verify that the patterns observed are not artifacts of keyword selection.
The findings suggest that Korean historiography has changed less by replacing its classificatory scaffolding than by reweighting agendas and reconstituting conceptual relationships within it. Dynasty- and polity-centered organizational categories remain structurally prominent throughout the period, indicating continuity in periodization and field organization. Within that stable frame, however, macro-level trajectories signal significant shifts in emphasis and conceptual framing. The relative decline of minjok after the 2000s is interpreted not as a simple retreat of the national category from Korean historiography, but as its recontextualization within new conceptual networks — a shift made visible through co-occurrence analysis rather than frequency counts alone. Keywords related to women increase markedly from the 1990s onward and consolidate into a stable, institutionalized research domain by the 2000s and 2010s, indicating the mainstreaming of gendered perspectives in Korean historical research.
In modern history research, a parallel process of conceptual reframing is observable. The period-defining label "Japanese rule" prominent in the 1970s–1990s gives way after 2000 to more analytical categories such as colony, Japanese colonial period, and East Asia, indicating a shift from periodization-oriented to analytically and comparatively oriented frameworks. The growth of East Asia as a research category further suggests a scaling of Korean historiography toward regional and transnational comparative frames. Topic clustering results situate these keyword-level observations within broader thematic structures, allowing the study to analyze not only which terms rise or fall, but how the relative weight of entire research clusters shifts across decades.
6. Conclusion
This paper represents a full-scale attempt to convert the BKHS into a large-scale bibliographic dataset and to analyze the temporal consciousness of Korean historiography at the macro level. The findings indicate that Korean historiography has not replaced its existing periodization and dynasty-centered classificatory framework, but has recalibrated the weight and relational configuration of particular concerns and conceptual vocabularies within it. The institutionalization of women's history, the recontextualization of minjok, the gradual emergence of environmental concerns, and the shift toward analytical and comparative frameworks in modern history research are among the most legible patterns this macro-level analysis surfaces.
The contribution of this study lies not in proposing new text mining techniques, but in reconstituting a foundational bibliographic resource as a dataset and enabling a scale of historiographical analysis that close reading-based trend reviews cannot achieve. By combining word-, relational-, and document-level analysis, the paper moves beyond simple frequency counting to examine how concepts emerge, persist, are recontextualized, and cluster over the long term. More broadly, as an empirical case involving a large non-English and non-Romanized scholarly dataset, the paper contributes to international Digital Humanities research by demonstrating that distant reading can be productively applied beyond the Western literary corpora for which it was originally developed, and that Digital Humanities can extend the research questions of historical scholarship rather than merely apply technology to existing problems.