Daejeon, July 27–31
This panel examines data-making as an interpretive practice that requires ongoing negotiation between humanistic and technical perspectives. Focusing on East Asian cultural materials, it explores how data formats function as epistemic frames, critically examines AI-assisted tools, and fosters productive dialogue between humanities scholars and technical practitioners in digital humanities contexts.
The fourteen contributions gathered here span provenance modelling, historical corpus construction, named-entity recognition for Literary Sinitic, computational pragmatics, sociolinguistic variation, and comparative political discourse. Taken together they treat modelling decisions — what counts as an entity, where a record ends, which silences a schema reproduces — not as preliminaries to interpretation but as interpretation itself.
Presenter(s): Min-gyu Shin, Korea Heritage Service; In-tae Ryu, Chonnam National University
Provenance information on Korean cultural heritage located overseas — the circumstances of its removal and movement, changes in ownership and custody, and records of trade and donation — survives only in fragments scattered across diverse sources. Although nearly four decades of state-led field surveys have produced a substantial body of research, most of the results have been published as linear texts, making systematic management and utilization difficult. Against this background, the Overseas Korean Cultural Heritage Foundation commissioned the Basic Study of Overseas Korean Cultural Heritage Provenance Data Modeling in 2022, and a three-year successive research project (2025–2027) is currently underway to deepen and develop its outcomes.
The most distinctive feature of the data model designed in this research lies in the Provenance Point class. Whereas Western provenance data models generally focus on tracing transfers of ownership, changes of location, and transaction histories, this model extends the concept of provenance beyond such boundaries to encompass the entire life cycle of an object, defining twenty-three types of provenance points ranging from creation and excavation to inheritance, trade, auction, plunder, donation, exhibition, conservation, and damage.
These provenance points are established as an independent class, around which classes such as Person, Group, Place, Event, Artifact, and Sources are interconnected in a provenance-centric structure, enabling the historical contexts surrounding an object to be represented within a single network. In addition, the model sets up Memory Institution (museums and libraries) and Heritage Market (auction houses and antique dealers) as separate classes to capture the particularities of the world of overseas Korean cultural heritage, while its generalizable architecture — bound to no specific object or collection — allows museums, libraries, and other heritage institutions to employ it efficiently for object history management, research, and investigation.
Presenter(s): Woongchul Shin, Hanbat National University
[Abstract to be supplied by the presenter.]
Presenter(s): Ju-hee Choi, Duksung Women’s University
[Abstract to be supplied by the presenter.]
Presenter(s): HyunJeong Lee, Korea University
[Abstract to be supplied by the presenter.]
Presenter(s): Young-in Kim, Korea University
[Abstract to be supplied by the presenter.]
Presenter(s): Eunkyo Suh, Yonsei University
Kaesong merchants (開城商人), also known as Songsang (松商), were among the most influential commercial groups in late Joseon Korea. Particularly during the late nineteenth century and the opening-port period, they constituted one of the last major indigenous commercial networks operating amid the expansion of foreign capital and rapidly changing commercial structures. Despite their historical significance, however, sources related to Kaesong merchants survive only in fragmentary and uneven forms. As a result, existing scholarship has largely relied on qualitative approaches centered on institutional history or regional commercial history. This study explores how the activities and commercial networks of late Joseon Kaesong merchants may be reconstructed and traced through digital humanities methodologies using fragmented historical sources.
The project focuses on reconstructing merchant networks through a comparative examination of scattered late Joseon materials alongside more detailed modern and colonial-era commercial records. Primary sources include merchant account books, records related to the Sagae Songdo bookkeeping system (四介松都置簿法), late Joseon administrative documents, colonial-era commercial surveys, and materials concerning ginseng (人蔘) distribution and trade. Given the fragmentary nature of late Joseon sources, modern commercial records are treated not simply as materials from a later historical period, but as reference points through which earlier commercial relationships and patterns of activity may be retrospectively traced.
This process raises several methodological questions. How can relationships among merchants, Songbang (松房), and agents (差人) be reconstructed from incomplete historical evidence? What kinds of interpretive choices are involved in transforming fragmentary textual records into nodes, routes, networks, or structured datasets? Furthermore, how might different modeling approaches produce varying representations of ginseng distribution, merchant networks, and regional commercial connectivity?
Rather than presenting a completed digital database, this presentation focuses on the preliminary and exploratory stage of reconstructing late Joseon Kaesong merchant networks through digital humanities methodologies. At the same time, the study seeks to critically examine the methodological limitations and interpretive risks involved in retrospectively tracing and modeling fragmented historical materials through later sources. In particular, it recognizes as significant concerns the temporal gap between late Joseon and modern or colonial-period materials, the imbalance and uneven survival of sources, and the possibility that later categories and concepts may be excessively projected onto premodern commercial structures.
Accordingly, this study approaches digital modeling not as a form of neutral technical representation, but as an interpretive practice shaped by the absence and incompleteness of historical records. By critically engaging with these issues, this presentation aims to develop a preliminary theoretical discussion on how incomplete East Asian commercial archives may be approached within digital humanities frameworks. In doing so, it argues that the case of Kaesong merchants can illuminate both the possibilities and the limitations of digital humanities approaches in reconstructing late Joseon commercial networks from fragmentary historical sources.
Presenter(s): Myungju Jeon, Korea University; Hwaryun Shin, Korea University
The purpose of this study is to analyze how the narrative of Bari-degi is portrayed in contemporary cultural content and to examine how modern audiences perceive Bari-degi as a character. The study focuses primarily on the webtoon Princess Bari and the musical Hongryeon, while also examining supplementary cases such as video games in which Bari-degi does not appear as the protagonist. Through this, the study aims to compare how the Bari-degi narrative is utilized in each form of content, particularly which elements — the Bari narrative or the narrative of the deceased — are more prominently featured.
The analysis focuses on the narrative representations in the webtoon and the musical in order to examine the prominence and function of the Bari-degi narrative. Additionally, by analyzing online audience reactions — such as comments on KakaoPage for the webtoon Princess Bari and Twitter responses to the musical Hongryeon — this study examines the relationship between the Bari-degi portrayed in the creative works and the Bari-degi perceived by the audience. Through this, the study aims to determine whether Bari-degi is perceived in contemporary cultural content as the protagonist of the Bari narrative, as a mediator of the narrative of the deceased, or as a character embodying a combination of both roles.
This study is significant in that it examines the contemporary adaptations of the Bari-degi narrative within the context of contemporary cultural content — such as webtoons, musicals, and games — to elucidate how classical narratives are reconstructed and received in today's media environment. Furthermore, by conducting both work analysis and audience response analysis, it will be possible to identify specifically the meaning of the Bari-degi character as perceived by the modern public.
Presenter(s): Jeffrey Tharsen, University of Chicago
This presentation will provide an overview of emerging trends in how digital methods are being applied to the study of Literary Sinitic texts. It will also briefly outline a prospectus regarding future developments.
Presenter(s): Woo-jeong Kim, Dankook University; David Hogue, Dankook University; Yunsoo Shin, Dankook University
Sinographic script and Literary Sinitic grammar have been used as the language of state officials and literati within the Koreanic communities of northeastern East Asia for some two millennia, likely beginning around the last century BCE. Kim Byung-joon traces the early transmission of Sinographic script to the establishment of the Lelang commandery (낙랑군 樂浪郡) on the Korean peninsula by the Sinitic Han (한 漢) imperial state in 108 BCE (Kim 2010: 29). In the most recent millennium, the tradition of using Literary Sinitic (referred to in Korean as hanmun 한문 漢文) as the written language of bureaucracy and creative literary composition was an eminently integral feature of the states and societies of both the Koryŏ (918–1392) and Chosŏn (1392–1920) periods.
In the ancient, medieval, and pre-modern eras, cultures centered around the production, exchange, and keeping of Literary Sinitic texts have thus deeply informed Korean political structures and practices, diplomacy, religion, philosophy, language, and historical memory; texts composed in Literary Sinitic make up a significant proportion of the cultural and mnemonic heritage of modern Korean communities. The ability of contemporary scholars, historians, and educated individuals to access and understand hanmun texts is fundamental to being able to accumulate, maintain, and create knowledge of human experience in Koreanic societies of the past and comprehend how Koreanic societies interacted with other societies in East Asia.
Hanmun literacy, which was the reserve of the gentry classes and therefore a rare skill in Korean society prior to the modern period, is just as rare today — if not more so — as it was traditionally. Apart from the very small proportion of individuals who have received training in reading Literary Sinitic, most in Korean society are unable to access the store of memory and culture embedded in hanmun texts directly; instead, those who are interested in hanmun texts depend on translations available in modern Korean, many of which are provided open-access by institutions such as the Institute for the Translation of Korean Classics (한국고전번역원 韓國古典翻譯院) through its Comprehensive DB for the Korean Classics (Institute for the Translation of Korean Classics 2025).
Concerningly, many scholars of Korean history are also among those who are hanmun illiterate; and interest in learning hanmun among the younger generations has waned since the middle of the last century, with the increasing dominance of the desire for English language learning and the centrality of Han'gŭl 한글, the script for native Korean, in ethno-nationalistic ideologies of state building and national community stewardship that have informed public policy-making since the end of the Japanese imperial occupation in 1945 (Kim 2025). The erosion of hanmun knowledge has continued through the conventionalization of digital technologies and the consolidation of the digital sphere as the primary format of literacy and knowledge transmission, since hanmun text production is for the most part not an ongoing practice, and readership in digital technologies — as in previous formats of text production and readership in the modern period — favors the most recently produced texts.
The upkeep of hanmun knowledge as a living practice in the digital age is also complicated by the fact that there are tens of thousands of distinct symbols, consisting of logographic and phonographic elements, that make up the Sinographic script and are placed together in manifold combinations to form words and set phrases. For Korean Sinographic script and usage, there are also diction and phrases that are unique to the Korean cultural and historical experience. Together, these features mean that the reading of Korean hanmun requires voluminous dictionary and glossary reference tools to support the comprehension of texts. Bringing hanmun reading culture into the digital realm means also integrating these sources into that experience.
Given these challenges, in order to promote the upkeep of hanmun knowledge and the reading of hanmun texts as a living practice, the Dankook University Institute for Han-Character Education Research (단국대학교 한문교육연구소 檀國大學校漢文教育研究所, IHER) has since 2022 been developing a Named Entity Recognition (NER) system for the annotation of Korean hanmun texts (Dankook University Institute for Han-Character Education Research 2025). Since 2016, IHER has been involved in a yearslong research project, Studies in Converting Traditional Literary Sinitic Texts into Knowledge Information Structures Using AI-Based Data Processing, funded by the National Research Foundation of Korea.
The system in its current form, the Vocabulary Matching Tool for Classic Hanmun Texts (한문고전 어휘 매칭툴 漢文古典語彙匹配工具), is a web-based program hosted on IHER's open-access InDiGo Cloud platform (InDiGo Cloud 2025). The Matching Tool allows users to paste a text-formatted passage into the user interface, and then presents a list explaining in modern Korean the place names, personal names, bureaucratic titles, and generic vocabulary that appear in the pasted passage.
The tool is programmed with some 660,000 dictionary and glossary entries — the current exact count is 661,918 — that have been extracted from twelve Korean-language reference sources for hanmun language and structured into machine-readable datasets (Table 1).
| No. | Korean | Sinography | English |
| 1 | 단국대 한한대사전 | 檀國大學校漢韓大辭典 | The Dankook University Comprehensive Hanmun-Korean Dictionary |
| 2 | 단국대 한국한자어사전 | 檀國大學校韓國漢字語辭典 | The Dankook University Dictionary of Korean Sinographic Script |
| 3 | 조선왕조실록 전문사전 | 『朝鮮王朝實錄』專門辭典 | “Veritable Records of the Joseon Dynasty” Specialized Dictionary of Terms |
| 4 | 조선조 관직 정보 | 朝鮮朝官職情報 | Information on Joseon Dynasty Official Titles |
| 5 | 장서각 소장 의궤 수록 반차도 복식 사전 | 藏書閣所藏儀軌收錄班次圖服飾辭典 | Dictionary of Clothing and Ornamentation Appearing in Procession Diagrams Contained in Joseon Dynasty Manuals of Protocol for State Rites Housed in the Jangseogak Archives |
| 6 | 조선조 문과급제자 | 朝鮮朝文科及第者 | Successful Candidates in the Joseon Dynasty Higher Civil Service Literary Examination |
| 7 | 조선조 사마시급제자 | 朝鮮朝司馬試及第者 | Successful Candidates in the Joseon Dynasty Lower Civil Service Literary Examination |
| 8 | 조선조 잡과급제자 | 朝鮮朝雜科及第者 | Successful Candidates in the Joseon Dynasty Civil Service Examination for Miscellaneous Areas of Expertise (e.g. Law, Foreign Affairs, and Medicine) |
| 9 | 역대인물정보 | 歷代人物情報 | Information on Historical Personages Through the Ages |
| 10 | 조선왕조실록 부가정보 인물 데이터 | 『朝鮮王朝實錄』附加情報・人物資料 | “Veritable Records of the Joseon Dynasty”: Information on Appendices and Data on Historical Personages |
| 11 | 고전용어 시소러스 | 古典用語類語辭典 | Thesaurus of Terms Used in Classic Texts |
| 12 | 편목색인정보 | 編目索引情報 | Information on Catalogue Concordances |
Table 1: Sources of the annotations and definitions of vocabulary terms in the database of the Hanmun Matching Tool.
The breadth and richness of the database that is the foundation of the Matching Tool makes the system an invaluable resource both for new learners of hanmun and for researchers who have an advanced knowledge of Literary Sinitic but who inevitably need to look up recondite idioms and terms occurring in traditional Korean Sinographic texts, such as names of bureaucratic offices or uncommon vocabulary appearing in poetic imagery.
Integrated into its annotating functions, the Matching Tool also performs segmentation analysis of text passages and parsing of individual segments into pre-structured vocabulary categories: personal names, place names, official titles, animals, plants, and so on. These functions are the first step in developing a large language model program that can automatically translate hanmun texts into modern Korean.
This paper will provide a detailed description of the functions of the Hanmun Matching Tool using live examples with passages of classic hanmun works. It will also include a functionality comparison with a comparable annotative system, MARKUS, to demonstrate the relative advantages of the Matching Tool for Sinographic texts of the Korean tradition.
Presenter(s): Minhyeok Kwon, International Research Center for Japanese Studies
This presentation reports on an ongoing data-structuring project conducted as part of the project Establishing the Digital History. The present study takes as one of its central tasks the extraction and structuring of knowledge from historical sources into forms amenable to computation and machine learning. More specifically, it may be characterized as a pattern analysis of predicate-centered word relationships drawn from Japanese-language chronological sources — a foundational step toward the broader goal of constructing a Japanese historical knowledge dictionary. This presentation sets out the full analytical workflow alongside concrete data examples, reports on the methodological decisions arising at each stage and their consequences for the structure of the resulting data, examines the distributional tendencies observable in the predicate–term patterns assembled to date, and considers the outstanding tasks that remain on the path to an automated extraction model.
The study's data collection is not restricted to a single historical field but encompasses chronological materials drawn from heterogeneous domains, including military history, international legal history, cultural history, and political history. This cross-domain design reflects a principled methodological commitment derived directly from the project's analytical objectives, rather than a matter of convenience or curatorial preference. Training a reliable automated extraction model for predicate–term co-occurrence patterns demands sufficiently diverse data that is not biased toward the idiomatic predicate usage of any single subfield. By analyzing in parallel sources that describe historically disparate types of events — military campaigns, treaty negotiations, railway construction, and the production of cultural works — the study is positioned to capture systematically the range of predicates operative across Japanese historical discourse and the full diversity of their contextual usage.
The core analytical procedure for pattern identification comprises five stages. First, chronological materials are collected from multiple historical fields and normalized to a common data format. Second, the predicate of each sentence is extracted; predicates are treated as the primary descriptors of historical actions, states, and relations, and function as the organizational axis of the entire data structure. Third, the terms bearing syntactic and semantic relation to each predicate are identified. Fourth, the relational role of each term with respect to its predicate is assigned, rendering explicit the co-relational structure between predicates and terms; particular attention is paid to the fact that the combinatorial relations of any given predicate or term may vary considerably across different source documents. Fifth, each term is classified within a structured categorization scheme; the principled construction of this system is itself a central methodological concern of the project, as the consistency of classification will directly determine the quality of the predicate–term patterns extracted in subsequent analysis.
The study's ultimate aim is to develop an automated historical knowledge extraction system powered by AI or machine learning. This presentation constitutes a preparatory stage toward that system, focusing on the systematic classification of terms appearing in chronological sources and the empirical identification of predicate–term co-occurrence patterns. The predicate–term pattern data accumulated through this process will serve as the training basis for an automated extraction model; the structured knowledge thereby produced is expected to contribute, in turn, to the development of a large language model underpinning a comprehensive historical knowledge dictionary.
Presenter(s): Yoon-young Jeon, Korea University
A persistent challenge in computational irony detection is the gap between classification performance metrics and theoretical understanding of what makes a text ironic. Most prior work treats classifier failure as a technical limitation: an engineering problem of insufficient data, suboptimal architectures, or inadequate features. The present study proposes an alternative framing. Model failure, when systematically analyzed, constitutes reverse empirical evidence for pragmatic theory. Specifically, the structural patterns of classification errors made by text-input models provide computational corroboration of the three-layer information structure predicted by the echoic attribution account of verbal irony (Wilson / Sperber 2012).
A Korean social media irony corpus (test set N = 422; non-irony = 210, irony = 212) was used to evaluate four classification models designed around distinct linguistic hypotheses: Model A (TF-IDF), implementing the lexical surface cue hypothesis; Model B (Sentence-BERT utterance-only embedding), implementing the distributional semantic representation hypothesis; Model C (SBERT with formal markers including quotation marks, emoji, and explicit irony signals), implementing the formal marker cue hypothesis; and Model D (SBERT with context–utterance relational features including cosine similarity, L1/L2 distance, and length ratio), implementing the contextual contrast hypothesis. All models used a linear-kernel SVM with five-fold group-based cross-validation to prevent context leakage.
AUC ranged from 0.574 (Model A) to 0.646 (Model D). Pairwise DeLong tests with Holm correction produced no statistically significant inter-model differences (A vs. D: uncorrected p = .032, corrected p = .192). No model achieved robust or theoretically sufficient discrimination across any feature configuration. This uniformly modest performance is itself the central finding: textual surface information across lexical, semantic, marker-based, and relational feature types does not stably encode what is required to distinguish ironic from non-ironic utterances. Irony, these results suggest, is not a property of text but an inferential achievement of the reader.
McNemar tests at threshold 0.5 revealed qualitatively distinct error structures between Model A and the SBERT family (all A-vs.-SBERT comparisons p = .001 to .004; intra-SBERT non-significant, B vs. D: p = .345). Model A collapsed to near-zero specificity (0.029), producing 204 false positives out of 210 non-ironic instances and systematically treating emotionally charged vocabulary as irony signals regardless of discourse function. This pattern demonstrates that affective expression is not an irony-specific cue but a general engagement resource, a distinction that surface-level and embedding-based models alike fail to encode.
The primary contribution is a linguistically grounded taxonomy of model failure showing that false negatives cluster around three structurally missing information layers. A three-type taxonomy of false negative errors was derived from linguistic coding of 15 false negative cases and 15 matched true positive cases within the clear-and-no-marker subset (N = 67) of Model D. This subset comprises instances rated as unambiguously ironic yet lacking all formal surface markers, constituting the theoretically purest diagnostic window into structural classifier failure.
Type 1, unmarked irony, comprises cases in which the speaker's dissociative attitude is not encoded in any surface-level linguistic form. Recognition depends entirely on inferring evaluative distance from discourse context — precisely the information layer that no text-input model can access by design. This aligns with echoic attribution theory, in which evaluative stance attribution precedes irony recognition. The classifier's failure here is not incidental but structurally guaranteed.
Type 2, context-dependent echo, comprises cases in which the echoic target resides outside the target text. The thought, norm, or prior utterance being ironically echoed must be reconstructed from the preceding conversational or social frame rather than read from the utterance itself. This failure type reveals the inherent limitation of utterance-level classification for a fundamentally discourse-level phenomenon.
Type 3, socio-cultural norm breach, comprises cases in which recognition requires culturally shared background knowledge — institutional realities, generational contradictions, and politically resonant discourse — that is structurally absent from distributional corpus statistics. This type is particularly pronounced in Korean, where culture-specific emotional orientations toward social injustice and resigned critique shape ironic interpretation in ways no current embedding captures.
The false positive structure reveals a complementary fourth pattern: sentiment over-generalization, in which direct criticism is misclassified as irony because strong affective vocabulary occupies overlapping semantic space across ironic and non-ironic utterances. This confirms that emotional intensity and ironic stance are orthogonal dimensions, a theoretical distinction that classification architectures built on surface co-occurrence cannot encode.
These three false negative error types map precisely onto the three information layers central to echoic attribution theory: dissociative attitude, echoic target, and discourse premises. The classification system fails exactly where pragmatic theory predicts the inference bottleneck. We propose this correspondence as the foundation of an error-diagnostic paradigm: rather than optimizing for performance benchmarks, the productive question is which pragmatic information layers are structurally unavailable to text-input processing, and what this structural inaccessibility reveals about the nature of irony comprehension itself.
Presenter(s): Chanhee Lee, Kyoto University
This study provides a mathematical account of the ranuki-kotoba phenomenon — the omission of -ar- in the potential form of Japanese verbs. Specifically, we investigate the selection mechanisms between the standard form -rareru and the innovative form -reru, and infer the posterior probabilities of these choices in specific linguistic contexts. Recent decades have seen a significant proliferation of the -reru form in modern Japanese, and within Japanese linguistics the integration theory is widely accepted as the primary framework to explain this shift. Integration theory posits that the phenomenon is driven by a functional necessity to differentiate the potential meaning from other ambiguous functions of the -rareru suffix, such as the passive, spontaneous, and honorific meanings. The present research complements this intuition by mathematically modeling these shifts as pragmatic Gricean inference and the pursuit of communicative clarity.
Central to this study is the dialogue between humanistic interpretation and computational implementation. From a digital humanities perspective, we acknowledge that while large-scale datasets provide extensive empirical coverage, they often fail to capture the nuanced, invisible psychological processes of language users. To bridge this gap, we employ the Rational Speech Act (RSA) framework, which models communication as a recursive reasoning process between a speaker and a listener. By formalizing human-centric intuition into a Bayesian RSA model, we aim to simulate the black box of linguistic decision-making on the basis of Bayesian statistics.
The methodology follows a two-step Bayesian approach designed to address the inherent tensions in corpus data. First, we deliberately selected the Corpus of Everyday Japanese Conversation (CEJC) because its substantial volume and rich socio-linguistic annotations provide an optimal foundation for analyzing nuanced linguistic variation. However, we recognize that such corpus data is inherently subject to data bias depending on register and collection context. Indeed, an analysis of approximately 400 cases highlights this contrast: for the verb kuru ('to come'), the written corpus BCCWJ yields zero instances of the innovative form, whereas the CEJC reveals a roughly 2:1 ratio between the standard -rareru and -reru. To account for these biases and ensure the reliability of our grounded references, we perform Bayesian logistic regression as a rigorous statistical verification of the data. This step allows us to estimate the posterior distributions of socio-linguistic factors — such as stem length, speaker age, and gender — and quantify the precise degree to which each factor influences the speaker's choice of ranuki-kotoba.
Second, unlike previous RSA models that treat costs as fixed constants, this study defines utterance cost as a probability distribution derived from these verified socio-linguistic factors. By integrating these distributions into the RSA framework and employing Markov Chain Monte Carlo methods, we infer the posterior predictive distribution of a speaker choosing ranuki-kotoba in specific linguistic contexts. Ultimately, this research showcases how language change can be modeled as a dynamic interaction between speaker rationality and context-dependent costs, offering a computational case study of Bayesian inference that balances empirical data with interpretive depth.
Presenter(s): Yatong Xie, Korea University
This study compares Korean loanwords related to K-pop in Chinese-speaking and Spanish-speaking online fan communities. It examines how Korean words are used and adapted in different linguistic and cultural contexts. The study focuses on differences in writing, meaning, and social function, and analyzes how writing systems, internet culture, and fandom influence the spread and localization of Korean loanwords.
The most common Korean expressions in English-speaking communities include terms of address such as oppa, everyday Korean words, and K-pop fandom expressions (Khedun-Burgoine / Kiaer 2022). Because Korean uses a different writing system from English, there is no single standard way to write Korean words in the Latin alphabet. As a result, many fans change spellings based on pronunciation, personal preference, or internet usage. For example, although the official romanization of 언니 is eonni, many fans use unnie because it is easier to read in English-speaking communities. Internet culture also influences these expressions: hwaiting, for instance, may appear online as 5ting.
By September 2021, the Oxford English Dictionary had officially added 26 Korean words, including K-pop-related expressions such as oppa, unni, aegyo, Hallyu, daebak, and fighting. This shows that Korean language and culture have influenced other languages through the global spread of Korean popular culture.
This study compares how Korean loanwords are adapted differently in Chinese and Spanish online communities. Chinese and Korean share many Sino-Korean words and a historical connection through Sinographic characters, while Spanish mainly adapts Korean words through phonetic spelling in the Latin alphabet. Because of these differences, Korean words may show different patterns of adaptation in the two languages.
For the methodology, this study uses online posts from X/Twitter and Xiaohongshu as a small exploratory corpus. Posts related to K-pop are collected and analyzed in order to examine commonly used Korean expressions in online fan discourse. The study focuses on how Korean words are written and used differently in Chinese and Spanish communities.
The analysis also uses the adaptation stages proposed by Fadic (2002), who divided the adaptation of foreign words in Spanish into several stages, including phonological adaptation, orthographic adaptation, and semantic adaptation. On the basis of this framework, the study compares the level of adaptation of Korean loanwords in Chinese and Spanish online communities. As a work in progress, this project aims to explain how Korean words spread and change in different online communities.
Presenter(s): Nahyeon Lee, Ewha Womans University; Chaeyeon Jeong, Korea University
This study compares how major political parties in South Korea and Japan name and position gender issues in election manifestos. It examines manifesto documents issued for the five most recent Korean National Assembly elections and Japanese House of Representatives elections, focusing on South Korea's People Power Party and Democratic Party of Korea, and Japan's Liberal Democratic Party and Constitutional Democratic Party. Election manifestos are treated as official statements of policy preference as well as strategic texts through which parties foreground, subordinate, or reframe contested issues under electoral competition. Gender issues are therefore analyzed not simply as policy items, but as political material that reveals how parties manage gender conflict, backlash, voter mobilization, and policy prioritization.
The study constructs a comparative corpus by segmenting twenty manifesto documents into tables of contents, major sections, subsections, policy items, and paragraphs. Each unit is assigned metadata, including election year, party, document type, section hierarchy, and whether the unit appears as a heading. Gender-related vocabulary in Korean and Japanese is then organized into conceptual categories, including terms such as 여성, 성평등, 성차별, 젠더폭력, 저출생, 돌봄, 女性活躍, 男女共同参画, ジェンダー平等, and SOGI. The analysis measures normalized frequency, dispersion across policy items, appearance in headings and tables of contents, and hierarchical location within each document. Keyword-in-context reading and basic co-occurrence analysis are used to identify the policy language repeatedly associated with gender-related terms. On the basis of these corpus-derived patterns, the study analyzes how gender issues are framed in relation to low fertility, care, labor, safety, welfare, economic growth, rights, diversity, and support for vulnerable groups.
The comparative design highlights how gender issues are politically organized in the electoral manifestos of the two countries. In the Korean case, the study examines whether gender-related claims are positioned within adjacent policy domains such as low fertility, care, safety, labor, and welfare rather than being named as an autonomous equality agenda. In the Japanese case, it examines how explicit categories such as 女性活躍, 男女共同参画, ジェンダー平等, and SOGI are linked to policy rationalities such as labor-force mobilization, demographic policy, corporate reform, political representation, and rights. By distinguishing the visibility of gender language from its egalitarian orientation, the study explains how manifestos can display gender-related vocabulary while assigning it divergent political meanings. It contributes to comparative politics, party politics, and gender studies by showing how gender agendas are named, displaced, and strategically reframed in East Asian electoral politics.