DH 2026

Daejeon, July 27–31

Thu, July 3009:00–10:30S040206-208
Long Paper

Small Datasets and Big Questions.Perspectives on the application of unseen species models to Old Swedish courtly literature

Katarzyna Anna Kapitan
École nationale des chartes - PSL, Paris, France · katarzyna.kapitan@chartes.psl.eu
Paola Peratello
École nationale des chartes - PSL, Paris, France · paola.peratello@chartes.psl.eu

This paper presents preliminary results of applying unseen species estimators to
Old Swedish courtly literature. The study suggests that the survival rates vary depending on the form and genre of works included in the dataset. The results are based on an open-access dataset manually created for the purpose of this research, which due to its limited scale opens the question of applicability of some statistical methods to small historical datasets and interpretative challenges associated with them. Using point estimates achieved with different unseen species estimators, it is possible to compare the survival rates of works in Old Swedish with other European languages (Kestemont et al. 2022). With around 78% survival for works, the Old Swedish tradition positions itself among the better-preserved ones, alongside Icelandic (77%) and German (79%). Our study, suggests, however, that the category of courtly literature is not homogeneous, and it appears that, for example, works in prose are better preserved than works in verse. At the same time, Old Swedish being a small historical dataset, with only 31 works preserved in 88 witnesses, produces results with huge confidence intervals. This brings interpretative challenges we wish to discuss with the community to inform the next stages of our project.Old Swedish courtly literature. The study suggests that the survival rates vary depending on the form and genre of works included in the dataset. The results are based on an open-access dataset manually created for the purpose of this research, which due to its limited scale opens the question of applicability of some statistical methods to small historical datasets and interpretative challenges associated with them. Using point estimates achieved with different unseen species estimators, it is possible to compare the survival rates of works in Old Swedish with other European languages (Kestemont et al. 2022). With around 78% survival for works, the Old Swedish tradition positions itself among the better-preserved ones, alongside Icelandic (77%) and German (79%). Our study, suggests, however, that the category of courtly literature is not homogeneous, and it appears that, for example, works in prose are better preserved than works in verse. At the same time, Old Swedish being a small historical dataset, with only 31 works preserved in 88 witnesses, produces results with huge confidence intervals. This brings interpretative challenges we wish to discuss with the community to inform the next stages of our project.

Keywords: Unseen Species Models; Old Swedish courtly literature; Computational Philology; Textual Survival and Loss; Manuscript Studies.

  • Introduction

    This work is part of the Digital Approaches to the Survival and Loss of Old Norse Romances project (ANR-23-CPJ1-0118-01), CJM, ENC-PSL. The dataset and the notebooks documenting our workflow are archived and freely available on Zenodo at https://zenodo.org/records/20057070.

Addressing the conference theme “Beyond Patterns: Interpretation with Small Data,” this paper examines the possibilities and limitations of applying statistical methods to small historical datasets. Focusing on unseen-species estimators (USEs), which have recently been used to approximate the loss of medieval literatures (Kestemont and Karsdorp 2019; Kestemont et al. 2022a-b; Macedo 2025; Guidi et al. 2025), we discuss interpretative challenges encountered in our study of medieval Swedish courtly literature. Based on preliminary results, the paper has two aims: to present the first version of our open-access dataset together with initial USE-based approximations of the survival rates of Old Swedish courtly works, and to reflect on the role of data analysis and visualisation in shaping heuristic interpretation. The paper opens discussion on questions such as: Can USEs be applied reliably to small datasets? What counts as “small” in historical research? Can dataset suitability be assessed in advance? And how should results be interpreted? While we do not attempt to resolve these issues, we address them in the light of our findings—findings that arise from a research design governed by the following questions:

  • What are the survival rates for works of Old Swedish courtly literature, depending on the approximation of richness by different models, and how do they compare with other European traditions?
  • What are the survival rates for witnesses preserving Old Swedish courtly literature, and how do they compare with other European traditions?
  • Are there differences in the preservation of Old Swedish works depending on form (prose vs verse), theme (romance vs pseudo-history), or origin (translated vs indigenous)?
  • Dataset

The linguistic, thematic, and geographic scope of the study is defined as follows. We focus on medieval Swedish literature written in the vernacular: Old Swedish (OSwe). As the predecessor of Modern Swedish, OSwe developed from Old East Norse and was used in Sweden roughly between 1225 and 1525, providing both geographical and chronological boundaries. Thematically, the objective was to construct a dataset of Swedish romances and heroic narratives comparable to those produced by (Kestemont et al. 2022a), enabling meaningful cross-linguistic analysis.

Due to the lack of comprehensive digital reference works for OSwe romances, or for Swedish manuscripts and their contents more generally, the dataset had to be compiled manually through a heuristic process drawing on multiple resources. In this respect, Faymonville’s recent study investigating the transmission of courtly literature in Sweden was indispensable and served as a de facto foundation for our work. (Faymonville 2023) focused on courtly literature broadly understood. In establishing her corpus, she considered five properties: vocabulary, formulaic expressions, form, broadly defined theme of ‘adventure’, and temporal and spatial settings. This resulted in the inclusion of texts belonging to various genres, such as roman d’antiquité (e.g. Konung Alexander; Ahlstrand 1885-1862), roman d’aventure (e.g. Herr Ivan; Lodén 2012) and roman idyllique (e.g. Flores och Blanzeflor; Lodén 2021), as well as pseudo-historical works (e.g. Aff Danmarkis konongom; Lorenzen / Jørgensen 1930).

Building on (Faymonville 2023) for corpus definition, and on (Kestemont et al. 2022) for data structure and methodology, we created a dataset of 31 works preserved in 88 witnesses. The dataset includes titles of works, witnesses, shelfmarks and repositories where the manuscripts containing these witnesses can be found. Information on genre, form, and witness production dates was also provided, using (Faymonville 2023) and online resources such as ALVIN, ARLIMA, Fornsvenska textbanken, handrit.is, manuscripta.se, and Kålund’s and Andersson-Schmitt and Hedlund’s catalogues.

  • Methodology

    The theoretical and statistical groundwork of the USE method, including its adaptation for surviving textual witnesses, are thoroughly addressed in (Kestemont / Karsdorp 2019; Kestemont et al. 2022a-b). The present abstract builds on this foundation.

As noted above, several studies have applied USEs to approximate the original richness of cultural data, particularly in relation to the survival and loss of literature. These methods estimate the number of lost (i.e., unobserved) works from abundance data representing the extant literary tradition: each work is treated as a species, and each textual witness as a sighting of that species. The distribution of low-frequency works allows us to estimate unobserved texts and reconstruct the original richness of a tradition using different estimators. From these estimates, we can infer survival rates for works. An extension of this approach also allows to estimate the minimum number of sightings required to observe each unobserved species at least once, which can in turn be used to infer survival rates for witnesses. Estimators such as Chao1 (Chao 1984; Chao / Jost 2015), improved Chao1 (Chiu et al. 2014), ACE (Chao / Lee 1992), Jackknife (Burnham / Overton 1979), and Egghe–Proot (Egghe and Proot 2007) are implemented in Copia, a Python package for diversity estimation in cultural data developed by (Kestemont et al. 2022b; Kestemont et al 2022a; Karsdorp et al. 2023), which was used in this study.

  • Analysis and Results

To answer RQs 1 and 2, we first estimated species richness using five estimators available in Copia and confirmed that Chao1 is the most conservative, giving the lower-bound point estimate of around 40 works (Table 1). We then used Chao1 to reproduce the results of (Kestemont et al. 2022b) to place the Swedish tradition in perspective, demonstrating that Swedish is among the better-preserved literatures in terms of work survival (78%) and the best preserved in terms of witness survival (22%) (Tables 2 and 3; Figures 1 and 2). It should be noted, however, that Swedish has exceptionally wide confidence intervals, which makes the interpretation of the results difficult. We return to this point later.

Chao1iChao1ACEJackknife

Egghe-Proot

est39.89741.89741.62743.00042.051
lCI27.35228.10032.73133.39822.529
uCI69.75273.47153.20052.60194.761
DutchFrenchIcelandicGerman

Irish

Swedish
est0.4920.5350.7720.7890.8100.776
lCI0.6760.6410.9090.9150.9141.125
uCI0.3330.4310.6370.6520.7040.445
DutchFrenchIcelandicGerman

Irish

Swedish
est0.0750.0540.1690.1450.1920.221
lCI0.1350.0740.2880.2450.3055.301
uCI0.0400.0380.1010.0840.1250.056
Figure Figure 1. Works survival by language with point estimates (dot), and the range confidence intervals up to 1.
Figure Figure 2. Witness survival by language with point estimates (dot), and the range confidence intervals up to 1.

To address RQ 3 and its three sub-questions, we examined three sub-datasets defined by theme (romances vs pseudo-histories), form (prose vs verse), and origin (translations vs indigenous works). The most pronounced differences in survival rates appear between prose (95%) and verse (48%) (Figure 4), followed by romances (90%) and pseudo-histories (60%) (Figure 5), and to a lesser extent translations (89%) and indigenous works (74%) (Figure 6). The confidence intervals for each category are very wide, as summarised in Table 4 and illustrated in Figures 3–5 which present the Kernel Density Estimation (KDE) distributions of the bootstrap estimators from which the intervals were calculated.

Survival of works (Chao1)Survival of documents (minsample)
estlCIuCIest

lCI

uCI
romance0.9050.4991.3870.4680.066-1.241
history0.6040.2471.3620.1350.031-0.939
prose0.9470.4911.5450.6410.091-0.761
verse0.4750.1991.1650.0750.0220.505
translations0.8930.4531.4450.4330.066-0.791
indigenous0.7370.3631.2650.2170.0730.706
Figure Figure 3. Works survival rate (romance vs. pseudo-history).
Figure Figure 3. Works survival rate (romance vs. pseudo-history). — Figure 4. Works survival rates (prose vs. verse). —

Figure 5. Works survival rates (translations vs. indigenous).
  • Discussion

Several points must be emphasised. First, the extremely wide confidence intervals, which occasionally produce survival rates above 1 or below 0, reflect the limited size of the Old Swedish dataset and indicate that we cannot draw precise conclusions. In principle, bootstrap CIs simulate repeated sampling from a large population, but in studies like ours, they function as a diagnostic tool to assess the stability of the estimator. As a result, our findings can only provide coarse directional insights, for example, that prose appears better preserved than verse, rather than exact survival rates.

The instability of the results also highlights the questions regarding the influence of decisions made at the level of data collection and modelling. In our study, works are treated as “stories,” ignoring multiple versions, even though some works exist in several redactions. Classification decisions further affect the outcomes: for instance, (Faymonville 2023), includes one redaction of the story of Barlaam and Josaphat (i.e. the younger redaction derived from Speculum historiale; Johanterwage 2009; Cordoni 2014), but excludes another redaction (i.e. the older one derived from Legenda aurea; Johanterwage 2009; Cordoni 2014) which does not meet the five criteria of courtly literature defined in her study. Consequently, in our study we have only one story, which de facto only corresponds to one version. Would treating versions separately increase the number of singletons and improve estimation stability? Should the version derived from Legenda aurea also be included?

These and other questions we plan to address at the next stage of our project, as presented here dataset is a starting point for future expansion and refinement. Our goal is to construct a unified East Norse corpus, combining Old Swedish and Old Danish with additional granularity on the level of different versions and their witnesses. We hope that this larger dataset will yield more robust results which will allow for meaningful comparisons with the West Norse (Old Icelandic and Old Norwegian) and other European traditions.

References
  1. Ahlstrand, Johan A (1855–1862): Konung Alexander: en medeltids dikt: från latinet vänd i
  2. svenska rim omkring år 1380. (Samlingar utgivna av Svenska fornskriftsällskapet. Serie 1, bd
  3. 12.) Stockholm: Norstedt.
  4. ALVIN, <https://www.alvin-portal.org/alvin/home.jsf?dswid=-3396>, [07.05.2026].
  5. Andersson-Schmitt, Margarete / Hedlund, Monica (1988-1991): Mittelalterliche
  6. Handschriften der Universitätsbibliothek Uppsala : Katalog über die C-Sammlung, 1, 2, 3, 4,
  7. Uppsala, Uppsala universitetsbibliotek.
  8. ARLIMA, <https://www.arlima.net/>, [07.05.2026].
  9. Burnham, Kenneth P. / Overton, Walter S. (1979): Robust Estimation of Population Size When Capture Probabilities Vary among Animals. DOI: 10.2307/1936861.
  10. Chao, Anne (1984): “Nonparametric Estimation of the Number of Classes in a Population”, in Scandinavian Journal of Statistics 11: 265–270.
  11. Chao, Anne / Jost, Lou (2015): Estimating Diversity and Entropy Profiles via Discovery Rates of New Species. DOI: 10.1111/2041-210X.12349.
  12. Chao, Anne / Shen-Ming Lee (1992): “Estimating the Number of Classes via Sample Coverage”, in Journal of the American Statistical Association 87: 210–17.
  13. Chiu, ChunHuo / Wang, Yi-Ting / Walther, Bruno A. / Chao, Anne (2014): “An Improved Nonparametric Lower Bound of Species Richness via a Modified Good–Turing Frequency Formula” in Biometrics. Journal of the International Biometric Society, ahead of print. DOI: 10.1111/BIOM.12200.
  14. Cordoni, Constanza (2014): Barlaam und Josaphat in der europäischen Literatur des Mittelalters, Berlin/Boston, De Gruyter. DOI: 10.1515/9783110341898.
  15. Egghe, Leo / Proot, Goran (2007): “The Estimation of the Number of Lost Multi-Copy Documents: A New Type of Informetrics Theory” in Journal of Informetrics 1, 4: 257–68. DOI: 10.1016/j.joi.2007.02.003.
  16. Faymonville, Louise (2023): Hövisk litteratur och förändringar i det fornsvenska textlandskapet, Stockholm, Stockholms universitet.
  17. Fornsvenska Textbanken, <https://www.nordlund.lu.se/texter/aldre-fornsvenska/>, [07.05.2026].
  18. Guidi, Émilie / Moins, Théo / Camps, Jean-Baptiste (2025): “Transmission and Survival of Iberian Patristic Texts (3rd-5th Centuries)”, in Computational Humanities Research 2025, ed. by Taylor Arnold, Margherita Fantoli, and Ruben Ros 3. Anthology of Computers and the Humanities: 589–607. DOI: 10.63744/WVZDLY7xI2fT.
  19. Handrit.is, <https://handrit.is/> [07.05.2026].
  20. Johanterwage, Vera (2009): “Kung Avennir i Barlaams ok Josaphats saga – en hövisk härskare?”, in Maria Arvidsson & Karl G. Johansson (red.). Barlaam i nord: legenden om Barlaam och Josaphat i den nordiska medeltidslitteraturen, Oslo: Novus: 75–97.
  21. Karsdorp, Folgert / Kestemont, Mike / De Koster, Margo (2023): “Dark Numbers:Modeling the Historical Vulnerability to Arrest in Brussels (1879-1880) Using Demographic Predictors”, Long paper at DHBenelux 2023. Brussels.
  22. Kestemont, Mike / Karsdorp, Folgert (2019): Het Atlantis van de Middelnederlandse Ridderepiek : Een Schatting van Het Tekstverlies Met Methodes Uit de Ecodiversiteit. DOI: 10.2143/SDL.61.3.3287540.
  23. Kestemont, Mike / Karsdorp, Folgert / de Bruijn, Elisabeth et al. (2022a): “Forgotten Books: Supplementary Materials (Data and Code) (v0.3.1-Cr)”, Zenodo. DOI: 10.5281/zenodo.5947206.
  24. Kestemont, Mike / Karsdorp, Folgert / de Bruijn, Elisabeth et al. (2022b): “Forgotten Books: The Application of Unseen Species Models to the Survival of Culture”, in Science 375, 6582: 765–69. DOI: 10.1126/science.abl7655.
  25. Kålund, Kristian (1889): Katalog over den Arnamagnæanske håndskriftsamling, 1, Copenhagen, Gyldendalske boghandel.
  26. Lodén, Sofia (2012): Le chevalier courtois à la rencontre de la Suède médiévale. Due Chevalier au lion à Herr Ivan. Stockholm, Stockholms universitet.
  27. Lorenzen, Marcus / Jørgensen, Ellen (1930): Middelalderlig historisk litteratur paa modersmaalet: indledning og supplement til M. Lorenzens Gammeldanske krøniker. Copenhagen, Samfund til Udgivelse af gammelnordisk Litteratur.
  28. Macedo, Carolina (2025): “Modeling the Invisible: Applying the Unseen Species Model to Chivalric Literature in the Iberian peninsula”, in Computational Humanities Research 2025, ed. by Taylor Arnold, Margherita Fantoli, and Ruben Ros. 3. Anthology of Computers and the Humanities: 1428–1437. DOI: 10.63744/qH01jSZULykB.
  29. Manuscripta.se, <https://www.manuscripta.se/>, [07.05.2026].