Daejeon, July 27–31
This study aims to apply digital humanities methodologies to the theme of the frontier (邊塞, Byeonsae) in Sino-Korean poetry (Han-si) included in the Hankuk Munjip Chonggan(Korean Literary Collections in Classical Chinese). By doing so, it seeks to investigate the distribution patterns and qualitative characteristics of frontier poetry, establish a binary classification sys-tem to determine whether poetry belongs to this genre, and quantitatively map the topo-graphical landscape of Korean frontier poetry.
Originally, the frontier is a spatial concept in Chinese Poetry referring to the rugged northern borderlands confronting foreign tribes. In the history of Chinese literature, frontier poetry can be categorized into a narrow sense (狹義), strictly referring to a specific group of works from the High Tang period, and a broader sense (廣義), encompassing any poetry that uses the frontier as a motif to express emotions. To define it simply, it is "Han-si that reflects frontier life or is related to frontier life" (Choi 1994). In this frontier poetry, the depiction of "virtual land-scapes" (虛景) is particularly noteworthy. This refers to a literary frontier conceptually con-structed in the author's study, using only the thematic conventions of frontier poetry, solemn sentiments, and established ancient allusions (典故), without the author actually visiting the border. The literary frontier suggests a multi-layered space conventionally adopted to project the poet's desire for merit or the sorrow of living far from home, transcending simple geo-graphical boundaries. Furthermore, this concept was similarly transmitted to Korean Han-si (Lim 2000; 2003).
Since the frontier is a relatively clear and distinct theme in Han-si, various approaches have been taken in its study. Among them, research based on statistical figures (Chan 1978; Lim 2021) and attempts to extract specific genres and related poetic words from large-scale texts by integrating machine learning and natural language processing technologies (He et al. 2022) are particularly noteworthy. The latter study is especially notable for its intention to ex-tract frontier-themed poems from the Quan Tangshi and compile a thesaurus. However, exist-ing research required certain improvements, as it did not adopt the language model-based classification methods that have recently emerged as a strong alternative, and the extraction details of poetic words were too complex to grasp the main points easily.
Therefore, to supplement the limitations of previous studies and quantitatively investigate the topographical landscape of Korean frontier poetry, this study conducts research in the follow-ing three stages.
First, we construct a "Frontier Poetry Gold Standard Data" by selecting and classifying ap-proximately 200 poems each from prominent frontier-style poets of the Chinese High Tang period and early to mid-Chosun period poets noted for their prominent use of frontier imagery. This is done through literature verification and expert cross-validation, using the two groups as control subjects. This aims to contrast and train on the archetypal features of the frontier gen-re and the development of Korean frontier poetry that adopted them. This data, marked up in XML format and segmented into poetic word units through a preprocessing pipeline by expert researchers, serves as the ground truth data providing an absolute standard for model training and evaluation.
Second, based on the preprocessed data, we design an automatic classification model utiliz-ing a core thesaurus that encompasses the combinatory patterns of frontier-specific noun ob-jects (`term`) and their co-occurring (Co-occurrence) surrounding words. For the classifi-cation methodology applied here, we will comparatively analyze the applicability of intuitive rule-based matching, traditional machine learning utilizing word frequency and vector space division, and language model-based classification methods that grasp high-dimensional con-texts. Based on this, the verified optimal model will be expanded and applied to the previously prepared Hankuk Munjip Chonggan text corpus.
Third, we analyze the error rate occurring during the classification process and introduce an ensemble voting mechanism of multiple models to minimize misclassification. This aims to overcome certain limitations in frontier theme classification, such as when frontier vocabulary is borrowed but the poem actually has a stronger tendency toward object-description (詠物), or when a poem of sorrow (哀傷) is misclassified as a frontier poem due to lexical similarities. Furthermore, we enhance the classification model's Precision, Recall, and F1-Score through a Human-in-the-Loop process for data where models show disagreement.
As described above, this study is significant in its attempt to interpret the thematic complexity of classical Han-si using computational resources. It is expected to provide a methodological foundation for future automatic thematic classification and semantic network analysis based on multi-class classification.