Daejeon, July 27–31
Freytag’s pyramid (Freytag 1900) is a theoretical framework that defines the structure of a narrative according to five phases usually called exposition, raising action, climax, falling action and denouement. This article focuses on narrative segmentation of literary and historiographical texts and their annotation with Freytag’s pyramid labels using large language models (LLMs). The topic is of interest from a theoretical, technical and pedagogical point of view. Recent work in this area has showed that multi-agent generative AI systems (Arif et al. 2025) using Freytag's pyramid structure to guide narrative generation via LLMs along with multimodal outputs (text-to-speech, video, music) can produce engaging educational settings for story co-creation and enhance children’s storytelling skills. Other studies that incorporated LLMs into narrative generation and comprehension pipelines and comparison with human performances have reported both promise and challenges in tasks such as narrative topic labelling and detection of elementary components of narrative discourse (Piper / Bagga 2024; Piper / Wu 2025), as well as AI story writing (Tian et al. 2024).
In digital humanities, LLMs’ evaluation represents a key aspect that needs to be further examined due to the very nature of problem-solving in this domain, which can often imply various degrees of uncertainty and subjectivity not easily measurable by clear-cut methods. For example, how can we assess LLMs’ output and reasoning in complex tasks such as narrative segmentation and annotation? To what extent the training, mathematical and probabilistic mechanisms inherent to these technologies can reflect a certain degree of comprehension of the story and problem to be solved?
Our assumption is that to answer this question a more complex evaluation system would be required to take into account multiple perspectives on the analysed processes. To test this assumption, we propose a meta-benchmark that considers three types of tasks performed by the models: (1) annotation; (2) evaluation; (3) rebuttal. A human agent can also be involved as a meta-evaluator in all the three tasks. The rationale of such a scenario is that by asking the models to solve a problem, evaluate the solutions of other models (and itself) to that problem and defend their own way of solving it may shed light not only on the capabilities and limits of these models in assuming various roles within the problem-solving space but also on how humans can interpret this type of mirror effect sometimes reflecting pieces of their own image. The proposed framework includes a set of data and an experimental protocol as described in the following sections.
The size of the dataset used for this project is intentionally small to ensure closer control and analysis of results and interpretation. Two types of texts were chosen for the experiments: (1) literary; (2) historiographical. The literary texts are short stories (8.000 to 14.000 characters) published in the 1910s-1920s and available via Project Gutenberg (Montgomery 1910; Daudet 1920; Chekhov 1921; Mansfield 1922). The historiographical texts are short excerpts (4.600 to 5.700 characters) from four 19-th century historical works and editions also available in the public domain (Michelet 1847; Tocqueville 1856; Burckhardt 1878; Ranke 1887).
Although the experiment language is English, with both original and translated texts, the selected works involved authors of different origin and gender to ensure a certain level of diversity. The choice of the historiographical texts was driven by White’s (1973: 7) concept of emplotment, as the way by which a “sequence of events fashioned into a story is gradually revealed to be a story of a particular kind”, and its application to the analysis of four major 19-th century historical works from which the fragments were extracted. It was therefore intended to compare the segmentation and annotation of short stories and test whether a certain plot structure, as hypothesised by White, can be detected for the historiographical selection too. While the study is not meant to discern between different types of emplotment in White’s sense (i.e., shaping historical events into narrative forms by using literary structures such as romance, comedy, tragedy and satire), a simpler annotation schema based on Freytag’s pyramid was chosen for the experiments as a way of identifying basic narrative units and trajectories of the rise and fall in narrative tension. It was assumed that story segmentation into narrative fragments and their annotation with Freytag labels by the models require a certain degree of understanding of the story arc and tension distribution across the segments, and conformity to a strict linear order of the annotated episodes.
The experiments used three LLM-based chatbots (ChatGPT 5.2. Thinking, Claude Opus 4.5, Gemini Pro3 Fast) and three open Hugging Face models running locally in GGUF format (mistral-7b-instruct-v0.2.Q4_K_M, llama-pro-8b-instruct.Q4_K_M, SOLAR-10.7B-Instruct-v1.0.Q4_K_M). The choice of the models was determined by the author’s access to the first two chatbots via personal and institutional subscriptions and the free accessibility of the others. Selecting both big and small, commercial and open models was part of the plan of ensuring variety and a comparison basis for the study.
Figure 1 shows the main phases in the meta-benchmarking workflow. The models receive the same prompts (Appendix 1) and the corresponding files as input, depending on the phase (annotation, evaluation, rebuttal). Outputs are evaluated by other models and the human agent (Appendix 2), but a different architecture with additional layers and hierarchies of evaluation both by human and AI can also be envisaged.
Figure 1. Meta-benchmark workflow for Freytag annotation using LLMs
The project is still ongoing and the abstract reports on preliminary results. A first round of tests has been performed so far with the AI chatbots and open LLMs for the annotation, evaluation and rebuttal related to the literary and historiographical dataset.
Table 1 shows the evaluation results for the annotation of the literary (lit) and historiographical (hist) texts, with the highest scores obtained by Claude (hist) and ChatGPT (lit) from the chatbot category, and Mistral (hist) and Llama (lit) from the open LLMs. For the historiographical annotation, Claude received the highest scores from Gemini for both criteria, conformity with prompt instructions (C1) and quality of annotation (C2), and lower scores from ChatGPT and itself. We illustrate the literary annotation, evaluation and rebuttal for Daudet’s short story. The main issue in Claude’s annotation, pointed out by ChatGPT and itself (Gemini assigned to it the maximum score, while the open LLMs either provided no numerical score or a numerical score with less specific justification), consisted of an early and diluted climax that spanned 3 segments, and the placement of the tension peak "To arms! To arms! The Prussians!“ in the denouement instead of climax. Interestingly, none of the models identified this sequence as the climax, but ChatGPT and Claude mentioned it in their evaluations (Appendix 3).
Table 1. Evaluation of the annotation task (literature, historiography) with degrees of criteria (C1, C2, C1+C2) fulfilment on a Likert scale 1 (very low) to 5 (very high) (runs 13-14.12.2025 and 19-22.03.2026)
| Evaluator | Criteria | Annotator | |||||||||||
| ChatGPT | Claude | Gemini | Llama | Mistral | SOLAR | ||||||||
| Hist | Lit | Hist | Lit | Hist | Lit | Hist | Lit | Hist | Lit | Hist | Lit | ||
| ChatGPT | C1 | 3.875 | 4.562 | 4 | 4.25 | 3.937 | 4.375 | 2.375 | 3.062 | 3.5 | 3.687 | 2.5 | 2.937 |
| C2 | 4.312 | 3.625 | 4.062 | 3.25 | 3.687 | 3.062 | 2 | 1.375 | 1.75 | 1.687 | 1.812 | 1.687 | |
| C1+C2 | 4.093 | 4.093 | 4.031 | 3.75 | 3.812 | 3.718 | 2.187 | 2.218 | 2.625 | 2.687 | 2.156 | 2.312 | |
| Claude | C1 | 3.625 | 4.5 | 4.562 | 4.687 | 4.062 | 4.5 | 1.687 | 3 | 3 | 4.062 | 1.812 | 2.562 |
| C2 | 3.937 | 3.812 | 4.437 | 4.562 | 4 | 4.375 | 1.937 | 1.375 | 1.687 | 1.875 | 1.812 | 1.5 | |
| C1+C2 | 3.781 | 4.156 | 4.5 | 4.625 | 4.031 | 4.437 | 1.812 | 2.187 | 2.343 | 2.968 | 1.812 | 2.031 | |
| Gemini | C1 | 3.812 | 4.937 | 4.875 | 4.875 | 4.875 | 5 | 2.5 | 3.437 | 3.562 | 4 | 3.25 | 2.812 |
| C2 | 4.562 | 4.5 | 4.875 | 4.812 | 4.812 | 4.812 | 2.25 | 1.937 | 2.75 | 2.375 | 2.25 | 1.937 | |
| C1+C2 | 4.187 | 4.718 | 4.875 | 4.843 | 4.843 | 4.906 | 2.375 | 2.687 | 3.156 | 3.187 | 2.75 | 2.375 | |
| Llama | C1 | 3.125 | 2.25 | 2.875 | 1.75 | 0 | 0 | 3 | 1 | 3.25 | 1.25 | 2.5 | 1 |
| C2 | 2.187 | 2 | 2.625 | 1.687 | 0 | 0 | 2.812 | 0.75 | 2.687 | 1.062 | 2.25 | 1 | |
| C1+C2 | 2.656 | 2.125 | 2.75 | 1.718 | 0 | 0 | 2.906 | 0.875 | 2.968 | 1.156 | 2.375 | 1 | |
| Mistral | C1 | 4.687 | 5 | 5 | 1.5 | 2.25 | 1.562 | 3.75 | 4 | 5 | 0.687 | 1.937 | 2.5 |
| C2 | 3.937 | 4.75 | 4.375 | 2.5 | 3.625 | 2.437 | 3.75 | 4.25 | 4.875 | 1.25 | 2.375 | 2.5 | |
| C1+C2 | 4.312 | 4.875 | 4.687 | 2 | 2.937 | 2 | 3.75 | 4.125 | 4.937 | 0.968 | 2.156 | 2.5 | |
| SOLAR | C1 | 3.375 | 3.5 | 2.562 | 4 | 2.75 | 4 | 3 | 3.187 | 3.25 | 2.5 | 2.187 | 1.937 |
| C2 | 2.312 | 3.75 | 2.812 | 4.062 | 2.5 | 2.375 | 2.812 | 2.875 | 4.162 | 3.062 | 3.187 | 2.437 | |
| C1+C2 | 2.843 | 3.625 | 2.687 | 4.031 | 2.625 | 3.187 | 2.906 | 3.031 | 3.706 | 2.781 | 2.687 | 2.187 | |
| Average annotation | C1 | 3.75 | 4.125 | 3.979 | 3.510 | 2.979 | 3.239 | 2.718 | 2.947 | 3.593 | 2.697 | 2.364 | 2.291 |
| C2 | 3.541 | 3.739 | 3.864 | 3.479 | 3.104 | 2.843 | 2.593 | 2.093 | 2.985 | 1.885 | 2.281 | 1.843 | |
| C1+C2 | 3.645 | 3.932 | 3.921 | 3.494 | 3.041 | 3.041 | 2.656 | 2.520 | 3.289 | 2.291 | 2.322 | 2.067 |
In general, the chatbots followed more closely the instructions and produced more structured annotations, while the open LLMs sometimes struggled with the required format, story coverage, number of segments, and the strict order of Freytag’s pyramid structure. For both categories, it was observed that although some aspects were not correctly defined in their own annotations, the models were occasionally able to identify these errors and properly comment on them in their evaluations or rebuttals, an intriguing behaviour that would need further investigation.
The article proposes a meta-benchmark that includes three phases, annotation, evaluation and rebuttal, for human (work in progress) and AI assessment of plot annotation with Freytag’s pyramid labels by LLM-based chatbots and open LLMs. It assumes that this type of framework may shed light on the reasoning capabilities of the models and the levels of understanding implied by complex tasks such as segmentation and plot annotation in literary and historiographical texts. While the framework can be applied to other research fields, the study focuses on story comprehension rather than broader human-like understanding, which requires extended testing and larger varieties of tasks and data.
| Annotation prompt | Evaluation prompt | Rebuttal prompt | Parameters (open LLMs) |
Divide the entire story below into 5-20 narrative segments. Annotate each segment with a label for narrative tension according to Freytag's pyramid: E (Exposition), R (Rising Action), C (Climax), F (Falling Action), D (Denouement). The segments must cover the entire story (not only a part) in linear order: E segments, R segments, C segments, F segments, D segments. Output a tab-separated table with 3 columns and 5-20 rows: Segment_number\tFreytag_label\tFirst_5_words\n Segment_number must be a sequential number starting from 1. Freytag_label must be one of E, R, C, F, D. First_5_words must be the exact first 5 whitespace-separated words of the segment copied verbatim. Now annotate this story: {story} | Critically assess the justness and penalise the mistakes of the Freytag annotation below corresponding to the story: {story}. Assign a Likert score 1 (very low) to 5 (very heigh) for the degree of fulfilment of each of the following criteria and a justification of your score. If you cannot complete the evaluation due to missing or insufficient information, or other reasons, please indicate so clearly. Output format: Evaluation of the annotation of {story_ref} by {annotator} Criteria | Score (1-5) | Justification (max 20 words) Criteria to evaluate: 1. Conformity with the instructions of the prompt: {prompt} 1.1. Column names 1.2. Number of segments in the specified range 1.3. Segment identification by first 5 words 1.4. Assignment of a Freytag label 2. Quality of annotation 2.1. Logic of the segmentation 2.2. Coverage of the whole story 2.3. Consistency of Freytag label assignment 2.4. Order of annotated segments according to Freytag's pyramid structure Now evaluate this Freytag annotation: {annotation} | Provide an assessment on a Likert scale 1 (very unfair) to 5 (very fair) and a rebuttal for the evaluation of your Freytag annotation for the story below. Output format: Rebuttal addressed to {evaluator} for the evaluation of my annotation of the story {story_ref} Prompt that provided the annotation instructions: {prompt} Annotated story: {story} Your Freytag annotation: {annotation} Now assess and respond to this evaluation: {evaluation} | temperature = 0.3 top_k = 40 top_p = 0.9 seed = 42 repeat_penalty = 1.15 repeat_last_n = 256 |
| Evaluation of LLMs’ annotation | Evaluation of LLMs’ evaluation | Evaluation of LLMs’ rebuttal |
| Same criteria as in evaluation prompt (Appendix 1) | 1. Conformity with prompt instructions 1.1. Addressing each criterium 1.2. Assigning a score (1-5) for each criterium 1.3. Providing a justification for each score 2. Quality of judgement 2.1. Correlation between score and criteria fulfilment 2.1. Correlation between score and justification 2.3. Validity of justification | 1. Conformity with prompt instructions 1.1. Providing a fairness score (1-5) for the evaluation 1.2. Providing a rebuttal 2. Quality of defending one’s point of view 2.1. Correlation between fairness score and justness of the evaluation 2.2. Validity of rebuttal |
Annotation (Claude Opus 4.5, ADaudet_TheSiegeOf Berlin_1920, run 14.12.2025) | Segment_number Freytag_label First_5_words 1 E We were walking up the 2 E You see those four closed 3 E I found the old cuirassier 4 R Meanwhile, the news of the 5 R "Very well, I will tell 6 R It was not an easy 7 R Unhappily, no matter how fast 8 R The first day of the 9 R After that day, our military 10 R Meanwhile, the siege was progressing—not 11 R As time went on and 12 C Another day, when his bed 13 C Unfortunately we were not sufficiently 14 C Poor old Jouve! No doubt 15 F For an instant he thought 16 D Suddenly the silence of the |
Evaluation (Claude Opus 4.5 by ChatGPT 5.2 Thinking, ADaudet_TheSiegeOf Berlin_1920, run 14.12.2025) | Evaluation of the annotation of **ADaudet_TheSiegeOfBerlin_1920.txt** by **claude_opus_4.5_ADaudet_TheSiegeOfBerlin_1920_20251412_1809_annot.txt** | Criteria | Score (1-5) | Justification (max 20 words) | 1. Conformity with the instructions of the prompt: {prompt} | 5 | Meets required format, range, valid labels, and linear ordering E→R→C→F→D. | | 1.1. Column names | 5 | Exact required headers: Segment_number, Freytag_label, First_5_words. | | 1.2. Number of segments in the specified range | 5 | 16 segments, within 5–20 requirement. | | 1.3. Segment identification by first 5 words | 4 | First-5-words strings match story passages, but true segment boundaries can’t be verified from the table alone. | | 1.4. Assignment of a Freytag label | 5 | All labels are valid (E/R/C/F/D) with one per segment. | | 2. Quality of annotation | 2 | Major misplacement of climax/denouement: the death scene is labeled D, weakening dramatic accuracy. | | 2.1. Logic of the segmentation | 3 | Segments follow event shifts, but three consecutive “C” segments blur a single turning-point moment. | 2.2. Coverage of the whole story | 4 | Starts with opening line and ends with final line; full inclusion between points is not provable here. | | 2.3. Consistency of Freytag label assignment | 2 | “To arms!” shout and collapse are the emotional peak, but placed in D not C. | | 2.4. Order of annotated segments according to Freytag's pyramid structure | 2 | Formal order holds, but functional order fails: climax occurs at the ending, not earlier “C” blocks. | |
Rebuttal (Claude Opus 4.5 to ChatGPT 5.2 Thinking, excerpts, ADaudet_TheSiegeOfBerlin_1920, run 14.12.2025) | **Rebuttal addressed to:** chatgpt_5.2_thinking **For the evaluation of my annotation of:** ADaudet_TheSiegeOfBerlin_1920.txt […] ## Fairness Assessment: 3/5 (Somewhat Fair) While the evaluation correctly identifies technical compliance with the prompt instructions, it applies an overly rigid interpretation of Freytag's pyramid that dismisses legitimate narratological alternatives. ## Rebuttal ### 1. The Climax Placement Dispute The evaluator's core criticism—that the death scene must be the climax rather than denouement—reflects one valid interpretation but not the only defensible one. **My rationale for the current structure:** In Freytag's original model, the **climax is the point of no return**, not necessarily the most emotionally intense moment. Segments 12-14 mark this turning point: - Segment 12: Colonel sees the National Guard (first crack in the illusion) - Segment 13: The family becomes careless; he likely overhears - Segment 14: He prepares for what he believes is the French victory parade By the time he appears on the balcony (Segment 15, Falling Action), **the tragedy is inevitable**. His death (Segment 16, Denouement) is the logical resolution of forces set in motion during the climax. **The evaluator's alternative:** Placing "To arms! The Prussians!" as climax treats it as the turning point. However, this moment doesn't _change_ anything—it merely _reveals_ what was already determined when the colonel decided to go to the balcony. […] ### Conclusion A fair evaluation would score "Quality of annotation" and "Consistency" at 3-4/5, acknowledging a defensible alternative interpretation rather than treating it as a major flaw. The current 2/5 scores reflect interpretive disagreement presented as objective error, which is inappropriate for a task involving literary analysis where multiple valid readings exist. |