Daejeon, July 27–31
In recent work in Computational Literary Studies, audiobook performances have increasingly been discussed as a promising source for reception-oriented data that complements text-based annotations (cf. Pethe et al. 2025; Stiemer et al. 2025; Schmidt et al. 2019). This line of research is based, among other things, on a long-standing assumption in speech science and literary studies that reading aloud in front of or to an audience constitutes an interpretative process in its very own right (cf. Brand 2021; Madison & Hamera 2006; Meyer-Kalkus 2019). From this perspective, the spoken realization of a prose text can be understood as a situated interpretation that renders aspects of narrative structure and semantics, particularly plot keyness, perceptible in performance. Prior computational work has predominantly approached plot keyness, understood as the degree to which individual textual elements or events stand out as relevant to the overall plot and its progression (cf. Ware & Farrell 2022), through textual features such as eventfulness, character involvement, or summary-based measures of plot importance. In contrast, audiobook recordings enable us to examine timing patterns that emerge from interpretive reading performance, most notably pauses and variations in reading speed. These timing patterns may reflect how readers emphasize and structure narrative content. We therefore ask: To what extent are pauses and reading speed markers of the readers' perception of plot keyness in prose texts?
Our analysis draws on a subset of GeAnProse, a yet unpublished data corpus of German prose texts, annotated with a range of narrative phenomena, including sentence-level plot keyness. Plot keyness is defined as a sentence’s relevance to the overall plot and is operationalised via a summary-based measure capturing the proportion of human-written summaries that reference a given sentence (cf. Hatzel et al. 2023). To relate narrative keyness to reception-oriented features, we use audiobook timing data for the same texts, comprising multiple audiobook recordings by professional and amateur readers. All recordings were processed using a forced-alignment pipeline that maps spoken performances to the textual sentences and yields token-level timing information. From these alignments, we derive two sentence-level temporal features: pause duration before a sentence and reading speed, normalized against a reference text-to-speech system.
Building on prior work (Stiemer et al. 2025) that conceptualises pauses in audiobook readings as semantically conditioned and interpretively motivated rather than purely technical phenomena, our workflow integrates these timing features with independently annotated plot keyness. By conducting all analyses at the sentence level, we enable an exploratory assessment of whether and to what extent temporal patterns in oral performance, specifically pausing behavior and reading speed, align with narrative relevance as modeled in the text.
Our analysis of the audiobook timing data shows systematic differences across reader groups and weak but consistent links to narrative keyness. Professional readers produce longer pauses after sentences than amateur readers, indicating a more pronounced segmentation of the text in performance. Across all readers, we observe a very small but statistically significant correlation between pause duration following a sentence and its plot keyness (r = .049, p < .05), suggesting that narratively relevant sentences tend to be followed by slightly longer pauses. By contrast, reading speed shows only a non-significant tendency to correlate with plot keyness, underscoring the limited sensitivity of tempo as a marker of narrative relevance.
A differentiation between reader groups further supports this interpretation, as professional and amateur readers exhibit nearly identical reading speeds when pauses are excluded, whereas professional readers produce systematically longer pauses. This pattern is also reflected in text-level averages (see Table 1): across all works for which both reader groups are available, professional readers consistently pause longer before sentences (e.g., 0.84 s vs. 0.60 s in the text Das Erdbeben in Chili; 0.80 s vs. 0.63 s in the text Die Verwandlung), while mean sentence reading times show no uniform direction across texts. In Das Erdbeben in Chili, professional readers read substantially more slowly on average than amateurs (10.68s vs. 9.80s per sentence), whereas in Die Verwandlung, the professional reading is slightly faster (7.93s vs. 8.54s).
| Work | Mean Pause Before Sentence (in Seconds) | Mean Sentence Reading Time (in Seconds) |
| Das Erdbeben in Chili - Amateur | 0.60 | 9.80 |
| Das Erdbeben in Chili - Professional | 0.84 | 10.68 |
| Der blonde Eckbert - Professional | 1.00 | 8.14 |
| Die Verwandlung - Amateur | 0.63 | 8.54 |
| Die Verwandlung - Professional | 0.80 | 7.93 |
| Krambambuli - Amateur | 0.72 | 7.10 |
Table 1: Mean pause duration before sentences and mean sentence reading time (in seconds) for selected works, differentiated by amateur and professional audiobook performances.
This text-specific variability reinforces the interpretation that reading speed primarily reflects individual and text-dependent performance styles rather than narrative keyness, while pausing behaviour emerges as the more stable and discriminative temporal feature. Taken together, these findings highlight limits to inferring narrative keyness directly from audiobook timing features, particularly when reading speed is considered in isolation, while at the same time supporting pauses as a low-threshold, reception-oriented indicator of how readers structure and weight narrative texts in performance.
The findings presented here are exploratory and reflect both the potential and the limitations of the underlying data. The corpus is small and text-specific, allowing for the identification of relevant effects but not for strong quantitative generalisation. Future work will therefore focus on expanding the dataset by adding further texts and more audiobook performances per text, enabling a clearer separation of text-specific, reader-specific, and idiosyncratic performance effects.