DH 2026

Daejeon, July 27–31

Thu, July 3015:20–16:20S089107
Short Paper

Interpreting Metaphors in Traditional Chinese Medicine with Artificial Intelligence: A Comparative Study of Four Models

Siyuan Peng
School of Information Management, Wuhan University · siyuanpeng@whu.edu.cn
Ping Wang
School of Information Management, Wuhan University; Key Laboratory of Archival Intelligent Development and Service,NAAC · wangping@whu.edu.cn
Zhicheng Xie
National Archives Administration of China · 349536647@qq.com
Yuan Cheng
School of Information Management, Wuhan University; Key Laboratory of Archival Intelligent Development and Service,NAAC · cynthia@whu.edu.cn

Traditional Chinese Medicine (TCM), a millennia-old treasure of Chinese civilization, encompasses practices such as acupuncture, herbal medicine, and qigong, offering a holistic paradigm for understanding human health (Guenier et al. 2025; Wang & Chen 2023). Recognized in the West as a form of complementary and alternative medicine, TCM preserves a rich knowledge base that can inform Western medicine (WM) (Lukman et al. 2007; Fu et al. 2021). This knowledge is primarily transmitted through ancient TCM books (Yang et al. 2022). Translating these works is therefore crucial for cross-cultural communication and integration between TCM and WM.

A salient feature of ancient TCM books is their pervasive use of TCM metaphorical expressions (TCMMEs) (Tang et al. 2025), which map concrete source domains onto abstract target domains, aiding understanding of concepts in pathology, physiology, and treatment (Shi & Yue 2024). For instance, The Yellow Emperor’s Classic of Medicine describes the “stomach as the sea of water and grain,” using the sea’s vast capacity to conceptualize the stomach’s containing functions. Because TCM and WM rest on fundamentally different epistemological and cultural frameworks (Guenier et al. 2025), translation of ancient TCM books cannot rely on direct linguistic transfer. Instead, it first requires accurate interpretation of the TCM metaphors, which is essential for enhancing TCM’s cross-cultural comprehensibility (Chen et al. 2025). Metaphor interpretation in TCM focuses on identifying the most relevant property mappings between source and target domains to reveal the literal meaning of TCMMEs (Wang et al. 2023).

However, interpreting TCM metaphors remains a significant challenge, requiring nuanced cultural and linguistic knowledge unique to TCM (Sun et al. 2024) and involving a highly labor-intensive process (Zayed et al. 2020). Concurrently, the rapid development of artificial intelligence (AI), particularly large language model–based systems, has fueled growing interest in their metaphor-interpretation abilities. Some studies suggest that AI possesses considerable metaphor-interpretation capabilities. For example, Zibin et al. (2025) found that models such as ChatGPT-4 and Google Gemini performed well in interpreting Classical Arabic metaphors, and Ichien et al. (2024) showed that GPT-4 generated explanations for novel literary metaphors that surpassed those of university students. Conversely, other studies highlight metaphor interpretation as a persistent challenge for AI. For example, Zibin et al. (2025) also reported difficulties in interpreting colloquial Jordanian Arabic metaphors, and Tong et al. (2024) using the Metaphor Understanding Challenge Dataset (MUNCH)—which spans academic writing, news, fiction, and dialogue—demonstrated that models such as LLaMA and GPT-3.5 failed to fully interpret certain metaphors.

Despite the growing attention to AI-based metaphor interpretation, its application to TCM metaphors remains largely unexplored, leaving uncertain the extent to which AI can support this task. This research gap is further compounded by the intrinsic diversity of TCM metaphors, which are typically classified into natural, social, and philosophical categories (Shi & Yue 2024), suggesting that model performance may vary across types. To address these critical yet under-researched issues, this study evaluates four state-of-the-art, publicly accessible AI models—including two international models (ChatGPT-4 and Gemini 3 Pro) and two Chinese models (DeepSeek-R1 and Kimi-K2)—for their ability to accurately interpret natural, social, and philosophical TCM metaphors.

Due to the lack of existing datasets in this field, this study is developing the Metaphor Interpretation in TCM Dataset (MITCMD) to assess the capability of AI in interpreting TCM metaphors. To ensure its representativeness and rigor, data are collected by conducting a subject search for the keyword “TCM metaphor” in the China National Knowledge Infrastructure (CNKI), which yields 235 relevant publications. As CNKI is access-restricted, it is unlikely that the selected AI models were previously trained on this material. Two experts specializing in both metaphor studies and TCM independently screen all publications to identify and extract sentences containing explicit explanations of TCMMEs. They record key elements such as the TCMMEs, source and target domains, mapped property pairs, and the corresponding ancient TCM books. If the experts do not agree with the extracted mapped property pairs, they provide their own explanations. All TCMMEs are classified into natural, social, or philosophical categories, with any discrepancies resolved through discussion. Examples from MITCMD are illustrated in Figure 1.

Figure 1. Examples from the MITCMD.

The four AI models are tasked with interpreting all TCMMEs in the MITCMD across the three metaphor categories. A unified prompt template is applied to each model. This template (1) positions the model as an expert in TCM metaphor interpretation and clearly specifies the task objective; (2) provides the TCMME together with its source and target domains; and (3) instructs the model to output its interpretation strictly in the format “source domain (projected property) → target domain (mapped property).” Example prompts and responses are shown in Figure 2.

Figure 2. Examples of AI model prompts and responses.

The annotated property pairs in the MITCMD serve as the gold standard. Two experts rate all AI-generated interpretations on a 5-point Likert scale. When uncertainties arise, they consult additional reference materials. Cohen’s kappa coefficient is computed to assess inter-rater reliability. The final score for each AI-generated interpretation is calculated as the average of the two expert ratings. To compare the overall performance of the four models in interpreting TCM metaphors, this study employs the Friedman test. Furthermore, the Aligned Rank Transform (ART) ANOVA is used to examine performance differences across different types of TCM metaphors.

In this short paper, we will present preliminary findings on the performance of four AI models in interpreting different types of TCM metaphors, followed by a focused discussion of both the potential and the challenges of using AI to support TCM metaphor interpretation. Building on these findings, we will propose strategies to enhance AI models’ capabilities in this domain, thereby facilitating the cross-cultural communication and application of TCM knowledge. This study contributes new empirical evidence to the ongoing discourse on AI-driven metaphor interpretation. Furthermore, by creating the MITCMD dataset, it addresses a notable gap in available resources and establishes a valuable data foundation for future studies on TCM metaphor interpretation.

References
  1. Chen, K., Hu, J., & Karabulatova, I. (2025). Translation adaptation of TCM (Traditional Chinese Medicine) terminology for speakers of other cultures: Features of cultural connotation. Philosophy, Ethics, and Humanities in Medicine, 20(1), 1–14.
  2. Fu, R., Li, J., Yu, H., Zhang, Y., Xu, Z., & Martin, C. (2021). The Yin and Yang of traditional Chinese and Western medicine. Medicinal Research Reviews, 41(6), 3182–3200.
  3. Guenier, A. W., Wang, B., Li, M., & Xing, M. (2025). How metaphors facilitate intercultural health communication: Insights from traditional Chinese medicine doctors in UK clinics. Journal of World Languages, 11(2), 319–342.
  4. Ichien, N., Stamenković, D., & Holyoak, K. (2024). Interpretation of novel literary metaphors by humans and GPT-4. In Proceedings of the Annual Meeting of the Cognitive Science Society (Vol. 46).
  5. Lukman, S., He, Y., & Hui, S. C. (2007). Computational methods for traditional Chinese medicine: A survey. Computer Methods and Programs in Biomedicine, 88(3), 283–294.
  6. Shi, Y., & Yue, F. (2024). Metaphorical thinking of traditional Chinese medicine and its features. World Journal of Traditional Chinese Medicine, 10(4), 535–547.
  7. Sun, Q., Karabulatova, I. S., Zou, J., & Kuo, C. (2024). Metaphorical terminology in ancient texts of traditional Chinese medicine: Problems of understanding and translation. Вестник Волгоградского государственного университета. Серия 2: Языкознание, 23(6), 141–157.
  8. Tang, J., Wu, N., Gao, F., Dai, C., Zhao, M., & Zhao, X. (2025). From metaphor to mechanism: How LLMs decode traditional Chinese medicine symbolic language for modern clinical relevance. arXiv.
  9. Tong, X., Choenni, R., Lewis, M., & Shutova, E. (2024). Metaphor understanding challenge dataset for LLMs. arXiv.
  10. Wang, F., & Chen, J. (2023). Translation studies of traditional Chinese medicine in China: Achievements and prospects. SAGE Open, 13(4), 21582440231204124.
  11. Wang, Z., Peng, S., Chen, J., Zhang, X., & Chen, H. (2023). ICAD-MI: Interdisciplinary concept association discovery from the perspective of metaphor interpretation. Knowledge-Based Systems, 275, 110695.
  12. Yang, F., Hou, J., Xing, C., Fu, X., Li, Q., Zhou, R., … Tao, X. (2022). Knowledge representation, acquisition and retrieval of ancient traditional Chinese medicine books based on knowledge element-semantic network model. Acquisition and Retrieval of Ancient Traditional Chinese Medicine Books Based on Knowledge Element-Semantic Network Model.
  13. Zayed, O., McCrae, J. P., & Buitelaar, P. (2020, May). Figure me out: A gold standard dataset for metaphor interpretation. In Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 5810–5819). Marseille, France.
  14. Zibin, A., Binhaidara, N., Al-Shahwan, H., & Yousef, H. (2025). Metaphor interpretation in Jordanian Arabic, Emirati Arabic and Classical Arabic: Artificial intelligence vs. humans. Humanities and Social Sciences Communications, 12(1), 1–12.