DH 2026

Daejeon, July 27–31

Thu, July 3013:40–15:10S064106
Short Paper

Large Language Models’ Capability Boundaries in Generating Xin Qiji’s Ci Poetry: A Humanistic Evaluation Study

Zhewei Liao
Rixin College, Tsinghua University, China, People's Republic of; Intelligent Laboratory for Chinese Traditional Culture, Tsinghua University, China, People's Republic of · liaozw22@mails.tsinghua.edu.cn
Shuang Xiao
School of Humanities, Tsinghua University, China, People's Republic of; Intelligent Laboratory for Chinese Traditional Culture, Tsinghua University, China, People's Republic of · xiaoshuang@tsinghua.edu.cn
Shengnan Sui
School of Humanities, Tsinghua University, China, People's Republic of; Intelligent Laboratory for Chinese Traditional Culture, Tsinghua University, China, People's Republic of · yiyuntianxin0128@163.com
Yafei Han
School of Humanities, Tsinghua University, China, People's Republic of; Intelligent Laboratory for Chinese Traditional Culture, Tsinghua University, China, People's Republic of · hyfhyf0913@163.com
Quanyu Lu
School of Humanities, Tsinghua University, China, People's Republic of; Intelligent Laboratory for Chinese Traditional Culture, Tsinghua University, China, People's Republic of · 1849252802@qq.com

Introduction

Since the release of ChatGPT in late 2022, large language models (LLMs) have rapidly entered domains of knowledge production and cultural creation. Yet for complex humanistic objects, their capability boundaries remain difficult to observe, partly because their generation process involves many opaque intermediate layers. Prior work cautions against equating LLMs’ statistical fluency with genuine understanding of language and cultural meaning (Bender et al., 2021). Xin Qiji’s (辛弃疾, 1140–1207) ci poetry, with its heroic yet often sombre style, is embedded in Song-dynasty literary culture and governed by strict generic conventions. It requires semantic continuity, prosodic control, historically situated diction, and the orchestration of refined and vernacular registers, making it a useful testbed for examining the gap between generative fluency and stylistic-cultural control. Digital humanities scholarship likewise shows that computational results depend on humanistic interpretation (Underwood, 2019).

This study uses prompt-based generation with mainstream Chinese and international LLMs to produce Xin Qiji-style imitations and evaluates them through quantitative metrics and expert assessment. By comparing lexical, thematic, semantic, formal, and affective performance, we treat ci imitation not merely as an LLM benchmark, but as a literary test case for identifying which aspects of Xin Qiji’s style resist statistical generation.

Research Design

Our corpus is drawn from Deng Guangming’s annotated Jiaxuan Ci: Chronological Edition with Commentary (《稼轩词编年笺注》), which we formatted and preprocessed to form the source dataset. We selected major LLMs, including DeepSeek, Kimi, and ChatGPT, and conducted two rounds of generation. The first round screened out systems with unstable outputs or severe genre deviations. For better-performing models, the second round added explicit requirements on line length and prosodic patterning to test whether external constraints could improve genre control. We then calculated quantitative indicators, screened candidate poems, and conducted expert scoring and close-reading-based analysis to cross-validate computational results against literary judgment.

Quantitative Analysis

We evaluate both model-level and single-poem performance with multiple indicators: LDA for thematic similarity, JSD for lexical divergence, BERT for semantic similarity, and EMD for overall distributional distance, later converted into an EMD-based similarity score. Overfitting is tested through source-corpus overlap, lexical overlap, n-gram repetition, and semantic redundancy. Similarity is calculated as follows:

Overall similarity = ωJSD×1-JSD+ωLDA×LDA+ωBERT×BERT+ωEMD×11+EMD

where each ω is the PCA-derived weight of the corresponding indicator.

Single-poem similarity = BERT semantic similarity.

Figure 1. PCA loading-matrix heatmap used to derive indicator weights.

To reduce subjective weighting, we use principal component analysis (PCA) to assign indicator weights. The input consists of 24 similarity values across six models and four dimensions. Indicators contributing more strongly to variance across models are treated as more informative and receive higher weights. Figure 1 shows the PCA loading matrix.

Because individual ci poems are short, JSD and LDA cannot extract enough lexical or thematic information at the single-poem level and therefore cannot reliably reflect similarity. BERT embeddings, by contrast, encode tokens into high-dimensional contextual vectors, here 768-dimensional, and carry richer semantic information. We therefore use BERT semantic similarity as the final single-poem indicator.

Figure 2. Overall similarity, mean single-poem similarity, and mean overfitting scores between generated imitations and Xin Qiji’s original ci poetry.

Figure 3. Single-poem similarity and overfitting scores of ci poems generated by different models.

The results show a clear asymmetry: generated imitations achieve relatively high thematic and semantic similarity to Xin Qiji’s originals, while lexical convergence remains weaker. JSD remains around 0.7, LDA around 0.8, and BERT similarity and EMD-based similarity both exceed 0.9. The maximum thematic and semantic values exceed 0.95, while lexical similarity based on 1 − JSD is mostly below 0.4. These results suggest that the imitations have reached a measurable level of differentiation, moving beyond surface resemblance toward semantic approximation, but fine-grained lexical and stylistic control remains weak. After optimisation, length and prosodic compliance improve markedly. Mean overfitting stays below 0.2, suggesting that overfitting to the source corpus has been effectively reduced (Figure 2). Single-poem similarity clusters around 0.79 and overfitting around 0.2, with low dispersion, indicating relatively stable quantitative performance (Figure 3).

Human Evaluation

On the basis of the quantitative indicators, we selected the ten highest-ranked imitations from each model. After removing segments with clear overfitting, fifty ci poems were retained for human scoring and qualitative analysis by six doctoral researchers specializing in Song-dynasty ci poetry and classical Chinese literature, who assessed semantic coherence, prosody, diction, emotional progression, and stylistic resemblance.

Human ratings reveal clear differences among models. Comparisons between Chinese and international models suggest persistent regional and cultural disparities. Although mean and median scores do not differ dramatically, ranges and variances vary considerably, indicating that output stability, rather than average performance alone, is a major limitation.

From an affective perspective, few imitations align well with Xin Qiji’s characteristic emotional structure. Many display abrupt scene-emotion shifts, missing transitions, or loosely assembled expressive fragments, lacking sustained and layered intensity. Semantically, generated poems often remain close to the originals, but linguistically they struggle to balance literary elegance with vernacular directness. Formally, they cannot consistently follow strict prosodic patterns or use Song-dynasty diction, especially in non-Chinese models.

Figure 4. Results of human evaluation scores for ci poetry generated by different models.

Conclusion

Taken together, the quantitative metrics and human evaluation show that LLMs can simulate culturally specific and highly regulated poetic genres to a certain degree, but their understanding is still moving from surface imitation toward deeper control. Current models approximate ci poetry through semantic modelling, yet remain weak in strict prosody, historically situated diction, and complex affective construction. The boundary revealed here is one of semantic plausibility but weak formal and emotional control. Layered emotion and distinctive expressive style remain central advantages of human literary creation. Non-Chinese models also underperform comparable Chinese models, highlighting regional and cultural differences in LLM performance and development.

By combining computational indicators with expert judgment, this study offers a computationally grounded and interpretable framework for evaluating Xin Qiji-style ci generation and examines the difference between AI and human creation at the level of concrete textual practice. Future research may refine the indicators, improve the scoring structure, and apply similarly difficult tasks to further explore LLM capability boundaries and human–AI relations. To support replication and further research, we will openly share the training materials, generated poems, quantitative data, and human evaluation scores (download link: https://drive.google.com/drive/folders/1XBO7b9jOXexqwSb1DBH9DoKCkrUMOiWw?usp=sharing).

References
  1. Bender, Emily M. / Gebru, Timnit / McMillan-Major, Angelina / Shmitchell, Shmargaret (2021): “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”, in: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21). New York: Association for Computing Machinery: 610–623. DOI: 10.1145/3442188.3445922.
  2. Chen, Huimin / Yi, Xiaoyuan / Sun, Maosong / Li, Wenhao / Yang, Cheng / Guo, Zhipeng (2019): “Sentiment-Controllable Chinese Poetry Generation”, in: Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI-19): 4925–4931. DOI: 10.24963/ijcai.2019/684.
  3. Sun, Maosong 孙茂松 (2020): “诗歌自动写作刍议”, in: 字人文 0: 32–38.
  4. Underwood, Ted (2019): Distant Horizons: Digital Evidence and Literary Change. Chicago: University of Chicago Press.
  5. Wang, Zhaopeng 王兆鹏 (2024): “古代文学研究的数据来源、指标和意义”, in: 北京大学学报(哲版)6: 78–87.
  6. Xia, Chengtao 夏承焘 (1959): “辛弃疾词论纲”, in: 学评论 3: 122–132.
  7. Xin, Qiji 辛弃疾 / Deng, Guangming 邓广铭 (annot.) (1993): 轩词编. Shanghai: 上海古籍出版社.
  8. Ye, Jiaying 叶嘉莹 (1987): “论辛弃疾词的艺术特色”, in: 文史哲 1: 44–54.