Daejeon, July 27–31
Traditional Chinese Medicine (TCM), a millennia-old treasure of Chinese civilization, encompasses practices such as acupuncture, herbal medicine, and qigong, offering a holistic paradigm for understanding human health (Guenier et al. 2025; Wang & Chen 2023). Recognized in the West as a form of complementary and alternative medicine, TCM preserves a rich knowledge base that can inform Western medicine (WM) (Lukman et al. 2007; Fu et al. 2021). This knowledge is primarily transmitted through ancient TCM books (Yang et al. 2022). Translating these works is therefore crucial for cross-cultural communication and integration between TCM and WM.
A salient feature of ancient TCM books is their pervasive use of TCM metaphorical expressions (TCMMEs) (Tang et al. 2025), which map concrete source domains onto abstract target domains, aiding understanding of concepts in pathology, physiology, and treatment (Shi & Yue 2024). For instance, The Yellow Emperor’s Classic of Medicine describes the “stomach as the sea of water and grain,” using the sea’s vast capacity to conceptualize the stomach’s containing functions. Because TCM and WM rest on fundamentally different epistemological and cultural frameworks (Guenier et al. 2025), translation of ancient TCM books cannot rely on direct linguistic transfer. Instead, it first requires accurate interpretation of the TCM metaphors, which is essential for enhancing TCM’s cross-cultural comprehensibility (Chen et al. 2025). Metaphor interpretation in TCM focuses on identifying the most relevant property mappings between source and target domains to reveal the literal meaning of TCMMEs (Wang et al. 2023).
However, interpreting TCM metaphors remains a significant challenge, requiring nuanced cultural and linguistic knowledge unique to TCM (Sun et al. 2024) and involving a highly labor-intensive process (Zayed et al. 2020). Concurrently, the rapid development of artificial intelligence (AI), particularly large language model–based systems, has fueled growing interest in their metaphor-interpretation abilities. Some studies suggest that AI possesses considerable metaphor-interpretation capabilities. For example, Zibin et al. (2025) found that models such as ChatGPT-4 and Google Gemini performed well in interpreting Classical Arabic metaphors, and Ichien et al. (2024) showed that GPT-4 generated explanations for novel literary metaphors that surpassed those of university students. Conversely, other studies highlight metaphor interpretation as a persistent challenge for AI. For example, Zibin et al. (2025) also reported difficulties in interpreting colloquial Jordanian Arabic metaphors, and Tong et al. (2024) using the Metaphor Understanding Challenge Dataset (MUNCH)—which spans academic writing, news, fiction, and dialogue—demonstrated that models such as LLaMA and GPT-3.5 failed to fully interpret certain metaphors.
Despite the growing attention to AI-based metaphor interpretation, its application to TCM metaphors remains largely unexplored, leaving uncertain the extent to which AI can support this task. This research gap is further compounded by the intrinsic diversity of TCM metaphors, which are typically classified into natural, social, and philosophical categories (Shi & Yue 2024), suggesting that model performance may vary across types. To address these critical yet under-researched issues, this study evaluates four state-of-the-art, publicly accessible AI models—including two international models (ChatGPT-4 and Gemini 3 Pro) and two Chinese models (DeepSeek-R1 and Kimi-K2)—for their ability to accurately interpret natural, social, and philosophical TCM metaphors.
Due to the lack of existing datasets in this field, this study is developing the Metaphor Interpretation in TCM Dataset (MITCMD) to assess the capability of AI in interpreting TCM metaphors. To ensure its representativeness and rigor, data are collected by conducting a subject search for the keyword “TCM metaphor” in the China National Knowledge Infrastructure (CNKI), which yields 235 relevant publications. As CNKI is access-restricted, it is unlikely that the selected AI models were previously trained on this material. Two experts specializing in both metaphor studies and TCM independently screen all publications to identify and extract sentences containing explicit explanations of TCMMEs. They record key elements such as the TCMMEs, source and target domains, mapped property pairs, and the corresponding ancient TCM books. If the experts do not agree with the extracted mapped property pairs, they provide their own explanations. All TCMMEs are classified into natural, social, or philosophical categories, with any discrepancies resolved through discussion. Examples from MITCMD are illustrated in Figure 1.
Figure 1. Examples from the MITCMD.
The four AI models are tasked with interpreting all TCMMEs in the MITCMD across the three metaphor categories. A unified prompt template is applied to each model. This template (1) positions the model as an expert in TCM metaphor interpretation and clearly specifies the task objective; (2) provides the TCMME together with its source and target domains; and (3) instructs the model to output its interpretation strictly in the format “source domain (projected property) → target domain (mapped property).” Example prompts and responses are shown in Figure 2.
Figure 2. Examples of AI model prompts and responses.
The annotated property pairs in the MITCMD serve as the gold standard. Two experts rate all AI-generated interpretations on a 5-point Likert scale. When uncertainties arise, they consult additional reference materials. Cohen’s kappa coefficient is computed to assess inter-rater reliability. The final score for each AI-generated interpretation is calculated as the average of the two expert ratings. To compare the overall performance of the four models in interpreting TCM metaphors, this study employs the Friedman test. Furthermore, the Aligned Rank Transform (ART) ANOVA is used to examine performance differences across different types of TCM metaphors.
In this short paper, we will present preliminary findings on the performance of four AI models in interpreting different types of TCM metaphors, followed by a focused discussion of both the potential and the challenges of using AI to support TCM metaphor interpretation. Building on these findings, we will propose strategies to enhance AI models’ capabilities in this domain, thereby facilitating the cross-cultural communication and application of TCM knowledge. This study contributes new empirical evidence to the ongoing discourse on AI-driven metaphor interpretation. Furthermore, by creating the MITCMD dataset, it addresses a notable gap in available resources and establishes a valuable data foundation for future studies on TCM metaphor interpretation.