Daejeon, July 27–31
The written legacy of Augustine, bishop of Hippo (354–430), is among the most influential and well-studied corpora in Western culture (Pollmann / Otten 2013). During the Middle Ages, the formative period of this culture, Augustine enjoyed the highest authority among the Church Fathers, interpreters of Scripture, and teachers of the Christian faith, but also served as a linking agent connecting the medieval world to Late Antiquity and, further back, the Classical past (Saak 2012; Boodts / Dupont 2018).
To a significant extent, Augustine’s immense influence was broadcast through a particular part of his oeuvre, his sermons (Dolbeau 2021). Whether records of actual preaching or texts composed to imitate or prepare such discourses, with their relatively concise, accessible, and multi-purpose format these texts played an important role in medieval education, devotion, and liturgy (Boodts / Schmidt 2022).
From Augustine, many hundreds of sermons survive in thousands of identified manuscript witnesses (Dolbeau 2021). However impressive this number may appear at first glance, it becomes puzzling upon closer examination. Dozens of other homiletic corpora of various sizes from the same epoch are known; some are by figures of comparable caliber, such as Ambrose of Milan—and yet they have left virtually no trace (Dupont 2018). This highlights possibly the most challenging feature of data related to cultural production in the past—its inherent incompleteness, due to constant sampling and sifting under the influence of uncountable historical and material factors. The data available to us result from collection and compilation processes we do not control nor even sufficiently understand. All of this brings to the fore—before any interpretation!—the question of the representativeness of the sample transmitted to us. Which fraction of Augustine’s preaching do we actually have?
It is this question that we address in this paper, making recourse to Unseen Species Models (USM) originating from ecology (Good 1953; Orlitsky et al. 2016), where they have been used to estimate the number of species in ecosystems that are invisible in available samples.
Seminally applied to cultural data as early as the 1970s to assess the number of words Shakespeare might have known (Efron / Thisted 1976), USMs have recently grown in popularity owing to the increasing availability of data to which they can be applied. After the inaugural and influential re-introduction by Kestemont and Karsdorp (2020) and Kestemont et al. (2022), applications range from patristic and medieval literature (Camps et al. 2025; Guidi et al. 2025) to literary studies more broadly (Martynenko 2023), from history (Wevers et al. 2022) and demographics (Karsdorp et al. 2024) to musicology (Moss et al. 2025). The recent publication of the package copia by Kestemont and Karsdorp will further facilitate the use of these models, contributing to a more systematic exploration of their usability and limitations when applied to cultural history data.
The main contributions of this paper are as follows:
Following Kestemont et al. (2022) and subsequent applications, we model the corpus of Augustine’s preaching as species abundance. Sermons are species, while their occurrences or manifestations in medieval manuscripts are observations. Let 𝑆 denote the total number of distinct sermons in the tradition, observed and unobserved, and let 𝑋𝑖 denote the number of manifestations of sermon 𝑖. The 𝑛 extant manifestations can then be expressed as
For any non-negative integer 𝑟, let 𝑓𝑟 be the number of sermons attested by exactly 𝑟 manifestations:
Thus 𝑓1 and 𝑓2 are the numbers of singletons and doubletons, while 𝑓0 is the unknown number of sermons not represented in the extant manuscript tradition. Total richness can therefore be expressed as
The task of USMs is to estimate 𝑓0. To achieve this, they introduce a probabilistic description of the transmission process. Each sermon 𝑖 is assumed to have an underlying, unknown probability 𝑝𝑖 of being observed in the current sample. This probability encapsulates all historical factors and circumstances of transmission—selection and copying of texts, transmission in collections, catastrophic losses, etc. The observed counts 𝑋𝑖 in Equation 1 can be understood as realizations of these probabilities. They need not be identifiable at the level of individual sermons; what matters is their effect on the shape of the observed sermon frequency distribution. In particular, traditions in which there are many texts with very low frequencies, would have a “long-tailed” text distribution, i.e., with many texts witnessed only once or twice. The core intuition behind USMs is that the presence of such rare but observed sermons contains information about the sermons that remain entirely unobserved.
As in Guidi et al. (2025), we estimate unseen richness using three estimators implemented in copia (Kestemont / Karsdorp 2025): Chao1 (Chao 1984), iChao1 (Chiu et al. 2014), and the fifth-order Jackknife (Burnham / Overton 1978). Chao1 estimates a lower bound from singletons and doubletons; iChao1 incorporates additional rare-frequency information; and the fifth-order Jackknife uses a broader range of low-frequency counts, 𝑓1–𝑓5, which is useful for a heterogeneous tradition such as ours. Finally, we use the Minimum Additional Sampling estimator to approximate the number of additional sermon manifestations that would suffice to observe all unseen texts.
Figure 1: Log scale sermons manifestation counts.
Table 1: Original Augustinian sermons in PASSIM dataset.
The study uses the core of the PASSIM dataset (Boodts et al. 2024)—an inventory of manuscript manifestations of authenticated Augustinian Sermones ad populum. A result of the work of several generations of scholars in France and Belgium, the inventory was from the mid-1970s until 2017 shaped by Luc De Coninck (KU Leuven), who enriched it mostly based on published manuscript catalogues, which is why the dataset reflects existing unevenness and biases in datafication of library holdings. The core PASSIM dataset describes almost 2700 medieval manuscripts produced from the early 6th to the first decades of the 16th century. Table 1 and Figure 1 provide basic statistics on the dataset.
Table 2 presents estimates of the total number of Augustinian sermons predicted by the three mentioned estimators. While estimates differ slightly, all suggest a very similar total number of distinct sermons (richness) converging at around 715 texts (or, if we subtract 599 observed sermons, approximately 115 unobserved texts [67–170 95% CI]). Discovering the unobserved texts, however, would require an unrealistically high number of witnesses, as suggested by the Minimum Additional Sampling estimator, which bases its calculation on the current frequency of rare sermons.
Table 2: Estimates of the total number of sermons (richness).
The results obtained can be interpreted on several levels. First—and optimistically—they may be taken as a modest reassurance regarding the quality of the data. Luc De Coninck’s inventory, made available through PASSIM, appears statistically very likely to cover the overwhelming majority of Augustine’s Sermones ad populum that entered the manuscript transmission. This emphasizes the exceptional nature of the discovery of completely new texts by François Dolbeau in the so-called Mainz codex (Dolbeau 1992).
Second—and less optimistically—these results, especially when confronted with historical evidence, confirm what scholars have long suspected: a substantial part of Augustine’s preaching is no longer accessible to us. Previous assessments of the total number of sermons preached by Augustine yield some 6,000 items, and while these assessments are considered unreliable and “peu significative” (Dolbeau 2021), what we know of Augustine’s career implies the preaching of several thousands of sermons. It has remained unclear how many of these thousands of sermons were ever recorded in writing and disseminated. Our estimations would suggest that the vast majority of Augustine’s sermons have, in fact, never entered the manuscript transmission as captured by Luc De Coninck’s inventory.
To further investigate this question, we can make use of a unique first-hand account of the early dissemination of Augustine’s sermons—a catalogue of his works attached to a biography written by his contemporary and friend Possidius of Calama (d. c. 437). An analysis of the structure of this so-called Indiculum suggests that this inventory may, at least in part, have mirrored the organization of the library of Hippo during Augustine’s lifetime (Dolbeau 1998). For this reason, it constitutes a valuable source of information about recorded sermons. Possidius’ Indiculum refers to a total of just under 300 sermons. Of these, only about one third have been matched with single or multiple identifiable sermons. The interpretation of Possidius’ list is methodologically complex, so it would be inaccurate to simply infer the existence of some further 200 sermons lost in our sample of the manuscript tradition. Still, Possidius suggests a number of unseen works greater than the consensus of the three estimators. This would be closer to commonsense- and historical-record-based expectations for more than three decades of a prominent preacher’s career. The tendency of these estimators to predict unseen works predominantly lower than historically expected has also been observed by Émilie Guidi in her analysis of Iberian Church Fathers, when compared with the Birth–Death model proposed by Camps et al. in (Guidi et al. 2025; Camps et al. 2025). In this study, we intentionally report only the results obtained with nonparametric estimators. We will dedicate a separate study to exploring the application of stochastic birth-death models, which are known to be highly sensitive to parameters describing historical reality (Camps et al. 2025; Guidi et al. 2025; Godreau et al. 2025).
This naturally leads us to reflect on the more fundamental problems entailed by applying USMs to cultural data. The assumptions and simplifications required for such a methodological bridge may violate the complexity and nuance of historical reality in hazardous ways. Two points are particularly noteworthy. The first is the assumption that observations or manifestations can be treated as independent. Although the problem does also exist in ecology, in the case of patristic sermons it is especially drastic. A single manuscript—containing a collection of manifestations of individual sermons—can contain a dozen rare texts, often not witnessed elsewhere. The PASSIM core dataset illustrates that ably. Among the rare sermons, only few occur independently, while most co-occur with at least one other rare sermon. It is not straightforward to determine whether such “interrelated” sermon observations should be treated as independent. Yet, doing so unavoidably inflates the estimated number of unseen sermons. The fact that such relatedness is a non-negligible factor emphasizes the second problematic point: unlike in ecology, where the elements of the population, namely species, can usually be defined quite clearly, this is not the case in transmission studies. Sermons were never “carved in stone”. Rather, they were constantly reworked, interpolated, excerpted, and recompiled. How should the resulting derivative works be treated? In this study, these complexities were normalized to the level of full sermons. Yet, the implications of this decision are to be explored more closely.
These methodological caveats as well as the discrepancy between estimates and the expectation supported by historical record should probably be taken as a reminder—especially vital in moments of data-driven neo-positivist enthusiasm—that the data available to us from centuries ago is the product of transmission processes of such complexity and variability that they—for the time being—may resist convincing modeling.