Daejeon, July 27–31
This paper presents preliminary results of applying unseen species estimators to
Old Swedish courtly literature. The study suggests that the survival rates vary depending on the form and genre of works included in the dataset. The results are based on an open-access dataset manually created for the purpose of this research, which due to its limited scale opens the question of applicability of some statistical methods to small historical datasets and interpretative challenges associated with them. Using point estimates achieved with different unseen species estimators, it is possible to compare the survival rates of works in Old Swedish with other European languages (Kestemont et al. 2022). With around 78% survival for works, the Old Swedish tradition positions itself among the better-preserved ones, alongside Icelandic (77%) and German (79%). Our study, suggests, however, that the category of courtly literature is not homogeneous, and it appears that, for example, works in prose are better preserved than works in verse. At the same time, Old Swedish being a small historical dataset, with only 31 works preserved in 88 witnesses, produces results with huge confidence intervals. This brings interpretative challenges we wish to discuss with the community to inform the next stages of our project.Old Swedish courtly literature. The study suggests that the survival rates vary depending on the form and genre of works included in the dataset. The results are based on an open-access dataset manually created for the purpose of this research, which due to its limited scale opens the question of applicability of some statistical methods to small historical datasets and interpretative challenges associated with them. Using point estimates achieved with different unseen species estimators, it is possible to compare the survival rates of works in Old Swedish with other European languages (Kestemont et al. 2022). With around 78% survival for works, the Old Swedish tradition positions itself among the better-preserved ones, alongside Icelandic (77%) and German (79%). Our study, suggests, however, that the category of courtly literature is not homogeneous, and it appears that, for example, works in prose are better preserved than works in verse. At the same time, Old Swedish being a small historical dataset, with only 31 works preserved in 88 witnesses, produces results with huge confidence intervals. This brings interpretative challenges we wish to discuss with the community to inform the next stages of our project.
Keywords: Unseen Species Models; Old Swedish courtly literature; Computational Philology; Textual Survival and Loss; Manuscript Studies.
This work is part of the Digital Approaches to the Survival and Loss of Old Norse Romances project (ANR-23-CPJ1-0118-01), CJM, ENC-PSL. The dataset and the notebooks documenting our workflow are archived and freely available on Zenodo at https://zenodo.org/records/20057070.
Addressing the conference theme “Beyond Patterns: Interpretation with Small Data,” this paper examines the possibilities and limitations of applying statistical methods to small historical datasets. Focusing on unseen-species estimators (USEs), which have recently been used to approximate the loss of medieval literatures (Kestemont and Karsdorp 2019; Kestemont et al. 2022a-b; Macedo 2025; Guidi et al. 2025), we discuss interpretative challenges encountered in our study of medieval Swedish courtly literature. Based on preliminary results, the paper has two aims: to present the first version of our open-access dataset together with initial USE-based approximations of the survival rates of Old Swedish courtly works, and to reflect on the role of data analysis and visualisation in shaping heuristic interpretation. The paper opens discussion on questions such as: Can USEs be applied reliably to small datasets? What counts as “small” in historical research? Can dataset suitability be assessed in advance? And how should results be interpreted? While we do not attempt to resolve these issues, we address them in the light of our findings—findings that arise from a research design governed by the following questions:
The linguistic, thematic, and geographic scope of the study is defined as follows. We focus on medieval Swedish literature written in the vernacular: Old Swedish (OSwe). As the predecessor of Modern Swedish, OSwe developed from Old East Norse and was used in Sweden roughly between 1225 and 1525, providing both geographical and chronological boundaries. Thematically, the objective was to construct a dataset of Swedish romances and heroic narratives comparable to those produced by (Kestemont et al. 2022a), enabling meaningful cross-linguistic analysis.
Due to the lack of comprehensive digital reference works for OSwe romances, or for Swedish manuscripts and their contents more generally, the dataset had to be compiled manually through a heuristic process drawing on multiple resources. In this respect, Faymonville’s recent study investigating the transmission of courtly literature in Sweden was indispensable and served as a de facto foundation for our work. (Faymonville 2023) focused on courtly literature broadly understood. In establishing her corpus, she considered five properties: vocabulary, formulaic expressions, form, broadly defined theme of ‘adventure’, and temporal and spatial settings. This resulted in the inclusion of texts belonging to various genres, such as roman d’antiquité (e.g. Konung Alexander; Ahlstrand 1885-1862), roman d’aventure (e.g. Herr Ivan; Lodén 2012) and roman idyllique (e.g. Flores och Blanzeflor; Lodén 2021), as well as pseudo-historical works (e.g. Aff Danmarkis konongom; Lorenzen / Jørgensen 1930).
Building on (Faymonville 2023) for corpus definition, and on (Kestemont et al. 2022) for data structure and methodology, we created a dataset of 31 works preserved in 88 witnesses. The dataset includes titles of works, witnesses, shelfmarks and repositories where the manuscripts containing these witnesses can be found. Information on genre, form, and witness production dates was also provided, using (Faymonville 2023) and online resources such as ALVIN, ARLIMA, Fornsvenska textbanken, handrit.is, manuscripta.se, and Kålund’s and Andersson-Schmitt and Hedlund’s catalogues.
The theoretical and statistical groundwork of the USE method, including its adaptation for surviving textual witnesses, are thoroughly addressed in (Kestemont / Karsdorp 2019; Kestemont et al. 2022a-b). The present abstract builds on this foundation.
As noted above, several studies have applied USEs to approximate the original richness of cultural data, particularly in relation to the survival and loss of literature. These methods estimate the number of lost (i.e., unobserved) works from abundance data representing the extant literary tradition: each work is treated as a species, and each textual witness as a sighting of that species. The distribution of low-frequency works allows us to estimate unobserved texts and reconstruct the original richness of a tradition using different estimators. From these estimates, we can infer survival rates for works. An extension of this approach also allows to estimate the minimum number of sightings required to observe each unobserved species at least once, which can in turn be used to infer survival rates for witnesses. Estimators such as Chao1 (Chao 1984; Chao / Jost 2015), improved Chao1 (Chiu et al. 2014), ACE (Chao / Lee 1992), Jackknife (Burnham / Overton 1979), and Egghe–Proot (Egghe and Proot 2007) are implemented in Copia, a Python package for diversity estimation in cultural data developed by (Kestemont et al. 2022b; Kestemont et al 2022a; Karsdorp et al. 2023), which was used in this study.
To answer RQs 1 and 2, we first estimated species richness using five estimators available in Copia and confirmed that Chao1 is the most conservative, giving the lower-bound point estimate of around 40 works (Table 1). We then used Chao1 to reproduce the results of (Kestemont et al. 2022b) to place the Swedish tradition in perspective, demonstrating that Swedish is among the better-preserved literatures in terms of work survival (78%) and the best preserved in terms of witness survival (22%) (Tables 2 and 3; Figures 1 and 2). It should be noted, however, that Swedish has exceptionally wide confidence intervals, which makes the interpretation of the results difficult. We return to this point later.
| Chao1 | iChao1 | ACE | Jackknife | Egghe-Proot | |
| est | 39.897 | 41.897 | 41.627 | 43.000 | 42.051 |
| lCI | 27.352 | 28.100 | 32.731 | 33.398 | 22.529 |
| uCI | 69.752 | 73.471 | 53.200 | 52.601 | 94.761 |
| Dutch | French | Icelandic | German | Irish | Swedish | |
| est | 0.492 | 0.535 | 0.772 | 0.789 | 0.810 | 0.776 |
| lCI | 0.676 | 0.641 | 0.909 | 0.915 | 0.914 | 1.125 |
| uCI | 0.333 | 0.431 | 0.637 | 0.652 | 0.704 | 0.445 |
| Dutch | French | Icelandic | German | Irish | Swedish | |
| est | 0.075 | 0.054 | 0.169 | 0.145 | 0.192 | 0.221 |
| lCI | 0.135 | 0.074 | 0.288 | 0.245 | 0.305 | 5.301 |
| uCI | 0.040 | 0.038 | 0.101 | 0.084 | 0.125 | 0.056 |
To address RQ 3 and its three sub-questions, we examined three sub-datasets defined by theme (romances vs pseudo-histories), form (prose vs verse), and origin (translations vs indigenous works). The most pronounced differences in survival rates appear between prose (95%) and verse (48%) (Figure 4), followed by romances (90%) and pseudo-histories (60%) (Figure 5), and to a lesser extent translations (89%) and indigenous works (74%) (Figure 6). The confidence intervals for each category are very wide, as summarised in Table 4 and illustrated in Figures 3–5 which present the Kernel Density Estimation (KDE) distributions of the bootstrap estimators from which the intervals were calculated.
| Survival of works (Chao1) | Survival of documents (minsample) | |||||
| est | lCI | uCI | est | lCI | uCI | |
| romance | 0.905 | 0.499 | 1.387 | 0.468 | 0.066 | -1.241 |
| history | 0.604 | 0.247 | 1.362 | 0.135 | 0.031 | -0.939 |
| prose | 0.947 | 0.491 | 1.545 | 0.641 | 0.091 | -0.761 |
| verse | 0.475 | 0.199 | 1.165 | 0.075 | 0.022 | 0.505 |
| translations | 0.893 | 0.453 | 1.445 | 0.433 | 0.066 | -0.791 |
| indigenous | 0.737 | 0.363 | 1.265 | 0.217 | 0.073 | 0.706 |
Several points must be emphasised. First, the extremely wide confidence intervals, which occasionally produce survival rates above 1 or below 0, reflect the limited size of the Old Swedish dataset and indicate that we cannot draw precise conclusions. In principle, bootstrap CIs simulate repeated sampling from a large population, but in studies like ours, they function as a diagnostic tool to assess the stability of the estimator. As a result, our findings can only provide coarse directional insights, for example, that prose appears better preserved than verse, rather than exact survival rates.
The instability of the results also highlights the questions regarding the influence of decisions made at the level of data collection and modelling. In our study, works are treated as “stories,” ignoring multiple versions, even though some works exist in several redactions. Classification decisions further affect the outcomes: for instance, (Faymonville 2023), includes one redaction of the story of Barlaam and Josaphat (i.e. the younger redaction derived from Speculum historiale; Johanterwage 2009; Cordoni 2014), but excludes another redaction (i.e. the older one derived from Legenda aurea; Johanterwage 2009; Cordoni 2014) which does not meet the five criteria of courtly literature defined in her study. Consequently, in our study we have only one story, which de facto only corresponds to one version. Would treating versions separately increase the number of singletons and improve estimation stability? Should the version derived from Legenda aurea also be included?
These and other questions we plan to address at the next stage of our project, as presented here dataset is a starting point for future expansion and refinement. Our goal is to construct a unified East Norse corpus, combining Old Swedish and Old Danish with additional granularity on the level of different versions and their witnesses. We hope that this larger dataset will yield more robust results which will allow for meaningful comparisons with the West Norse (Old Icelandic and Old Norwegian) and other European traditions.