Daejeon, July 27–31
Generative AI systems are increasingly used to support research on premodern texts, including tasks such as segmentation, annotation, and discourse-level analysis; however, their performance on classical languages remains unclear. Classical Chinese presents particular challenges for large language models (LLMs). Its semantics often diverge from modern usage, its syntax is compact and elliptical, and its discourse structure relies heavily on implicit rhetorical organization (Che, 2019). Since most LLMs are trained primarily on modern vernacular Chinese and English, it is not yet known whether these models can meaningfully capture the macro-level rhetorical patterns traditionally used by commentators to interpret classical texts. This study addresses this gap by evaluating how well contemporary LLMs align with established chapter-structure (章法) interpretations in Shiji, a foundational historiographical work whose narrative techniques have been extensively analyzed by classical commentators.
These are operationalized into a structured annotation scheme that maps traditional commentary categories onto a reduced and unified set of discourse function labels (OPENING, CONTINUATION, INTRODUCTION, ENTRY, TURN, SHIFT, FORESHADOWING, RESPONSE, SUSPENSION), further grouped into three higher-level functional categories: linear progression, transition, and discourse relation. To create a ground truth for evaluation, the study draws on Pu Qilong’s Guwen meiquan (《古文眉詮》), a Qing dynasty (Qianlong period, 1741–1744) annotated prose anthology compiled under Pu Qilong with contributions from Cheng Zhong and Fang Maofu, from which commentary-aligned annotations are extracted and systematically mapped to the proposed label set to allow comparison with machine-generated labels.
The experiment compares two leading LLMs—GPT-5.4 (Thinking mode) and Gemini 3 (Thinking mode)—across three prompting strategies: zero-shot, few-shot, and system-prompt with explicit guidelines. An annotation guideline was constructed to define each 章法 term with clear descriptions and examples. In the few-shot setting, each model was given several sample paragraphs with human-produced annotations following this guideline. The models were then asked to annotate the full ten-chapter corpus. Their output was compared with the ground truth using quantitative metrics, including macro F1, micro F1, Cohen’s kappa, label-level F1, and qualitative error analysis. Rather than testing predefined expectations, the analysis focuses on empirically observed patterns of agreement and systematic divergence between human annotations and model predictions. This approach builds on recent research evaluating LLMs’ discourse-level performance and coherence (de Wynter et al., 2023; Huber & Carenini, 2022).
Results show that explicit system-prompt guidelines improve model alignment with human chapter-structure annotations more reliably than few-shot examples. However, overall performance remains moderate, with best Macro F1 around 0.46 and Cohen’s kappa around 0.36. This indicates that current LLMs can partially recognize explicit and linearly signaled discourse functions but continue to struggle with structurally implicit and cross-segment rhetorical relations. A hierarchical analysis further demonstrates that large language models tend to collapse fine-grained discourse distinctions into a limited set of dominant categories, particularly favoring linear continuation structures. This tendency is especially pronounced in cases involving implicit discourse relations, where functions such as foreshadowing, suspension, and introduction are frequently misclassified as continuation. These patterns indicate a systematic limitation in LLMs’ ability to capture non-linear and cross-segment rhetorical structures in classical Chinese texts.
This study contributes to digital humanities in three ways. First, it offers one of the earliest systematic evaluations of LLM discourse-level semantic competence on premodern Chinese texts, demonstrating how classical rhetorical traditions can serve as a meaningful benchmark for AI evaluation. Second, it integrates traditional chapter-structure theory with contemporary annotation and alignment methodologies, creating a cross-cultural framework for discourse semantics. Third, the work expands the linguistic and cultural diversity of DH research by foregrounding non-Western interpretive systems, showing how historical commentary traditions can inform digital analysis. This aligns with broader calls for greater diversity and critical perspectives in LLM-based discourse studies (Gillings, 2024). By linking generative AI evaluation with classical Chinese textual scholarship, this study highlights not only the potential but also the structural limits of LLM-assisted research in humanities contexts where implicit discourse knowledge, cross-segment relations, and commentary-based interpretation are essential.
References