Daejeon, July 27–31
The emergence of machine learning (ML) methods and the widespread use of neural networks for the restoration and analysis of classical Greek and Roman inscriptions (Assael et. al. 2022, 2025) is changing the significance of existing epigraphical databases (e.g. EDCS, EDH, PHI). While individual use by scholars has been the main focus to date, these databases now provide datasets that are retrieved to be processed for further tasks in the field of machine learning or statistical analysis (Heřmánková et. al. 2021). Existing Bias in epigraphic datasets is therefore not only a methodological concern within epigraphy; it is a direct constraint on what machine learning, statistical analysis, and "distant reading" approaches can reliably infer from inscriptions. Most computational methods assume that the observed data are a sufficiently representative sample of a broader population, or at least that deviations from representativeness are known, measurable, and stable across the variables used for modelling. Epigraphic corpora rarely meet these assumptions. As a result, patterns detected by classifiers, embeddings, clustering, or distance-based retrieval can reflect the structure of collection and transmission rather than the structure of historical practice. This matters increasingly because computational approaches are now used not only for convenience (search and indexing) but for interpretive tasks: attribution, dating support, regional comparanda, genre detection, reconstruction assistance, and quantitative historical arguments based on distributional patterns.
Other historical disciplines have approached comparable issues by formalizing source criticism into their data practices: they treat representativeness as an empirical question, preserve uncertainty explicitly, and document how evidence is transformed by editorial and technical intervention (Hugget 2020). Epigraphy faces the same challenge, but the bias layers are often more tightly coupled to material survival, discovery contexts, reconstruction proposals, dating methods, and editorial conventions. As computational methods become routine, these implicit filters increasingly act as hidden variables in statistical comparisons and ML models. The relevant lesson is that bias is best handled as a structural constraint on inference rather than as noise that can be ignored or "cleaned away." For epigraphy, this means making bias and uncertainty visible enough that future analyses can choose appropriate methods, interpret results cautiously, and avoid treating dataset structure as historical structure.
This discussion is framed within the context of our ongoing project, which seeks to reconstruct worn inscriptions with optical measurements, whereby multiple data sets are created for further use. This is distinctive in that it does not merely aggregate conventional epigraphic data, as is customary in traditional corpora (Cayless et. al. 2009). Its primary function is to store advanced imaging data, including Reflectance Transformation Imaging, ultraviolet and infrared imagery, along with extensive metadata (e.g. Diez et. al. 2025). The creation of such a multifaceted dataset emphasises the urgency of this discourse. As sophisticated ML methods are introduced for dating, restoration and character recognition, datasets serve as the ground truth. Should there be any compromise of the integrity of this foundation due to unacknowledged biases, it is inevitable that the resulting computational models will produce skewed historical inferences. To this end, our project develops a documentation framework that records provenance, imaging parameters, and editorial decisions as structured metadata at each stage of the pipeline, enabling downstream computational workflows to query and account for known sources of uncertainty rather than treating all data points as equally reliable.
To address this, we must deconstruct the cumulative layers of bias introduced at every stage of an inscription's transmission. The first three layers belong to the analog domain, though their impact on digital methods remains under-discussed. The first layer is survivorship: only a fraction of ancient inscriptions were inscribed on durable media, fewer survived until modern times, and fewer were recovered. This constraint has long been recognized in classical scholarship (Gerstinger 1948). A second layer is findspot and topic bias. Epigraphic corpora are disproportionately shaped by highly visible sites and discovery contexts (e.g. major sanctuaries and urban centres such as Delphi and Athens). This affects the apparent geographic density of epigraphic habit, especially the distribution of formulae used for comparative work. While the present study is grounded in Greek and Roman material, these geographic and institutional asymmetries affect epigraphic corpora across traditions, and the layered bias framework proposed here is intended to be transferable beyond the classical Mediterranean. A third layer is editorial habit. Even when objects are available, editorial decisions filter what enters the published record and how it is represented. These decisions are transparent to epigraphers but can be opaque to computational workflows that treat an edition as an uniform text stream. For statistical analyses, inconsistent treatments of lacunae and restorations alter counts, co-occurrence patterns, and distance measures.
While these first three layers are widely acknowledged within the epigraphic community (Bodel 2023, Horster et. al. 2025, MacMullen 1982), digitization introduces a fourth layer of bias and inherit earlier ones. Encoding standards play a pivotal role in this regard. Corpora following EpiDoc are capable of preserving granular information about damage, expansion, restorations and editorial uncertainty, while plain-text transcriptions may collapse these distinctions. It is therefore possible for a model that has been trained across a variety of encoding regimes to interpret markup practices as substantive differences. Digitisation involves setting priorities when there are limited resources. This means deciding what to digitise first, what not to digitise, and when to only digitise part of something. The encoder can create inconsistencies, especially in data formats like dates, and can lead to major distortions further down the line. Expressions such as "before 200 BCE," "mid-2nd century BCE," or a "from–to" range encode different levels of uncertainty; if they are normalized naively, temporal modelling and diachronic distance calculations can become artefacts of conversion rules rather than reflections of historical chronology.
A fifth and final bias layer emerges at the point of access and reuse. Not all digitized inscriptions are openly available; restrictions range from paywalls to technical barriers (limited APIs, non-exportable interfaces, legal uncertainty). Researchers may resort to scraping, which introduces missing fields, parsing errors, and uneven extraction quality. This may cause an uncontrolled "data processing bias" that becomes embedded when datasets are redistributed and cited. Once such derived datasets circulate, errors and distortions can be difficult to trace back to their source, undermining reproducibility and creating feedback loops in which the most accessible corpora dominate training data and, consequently, computational conclusions.
In conclusion, the journey from the stone to the dataset is characterized by the accumulation and amplification of bias at multiple operational steps. As we transition into an era of data-driven epigraphy, we cannot afford to treat datasets as objective reflections of antiquity. This presentation advocates for the establishment of overarching guidelines and the rigorous application of FAIR and Open Data principles (Wilkinson et al. 2016, Heřmánková et al. 2022). Sustainable digital epigraphy requires a transparent methodological framework that acknowledges and documents these biases, ensuring that the datasets we build today remain valid foundations for the computational analysis of the future.
This research was funded by the Steiermärkische Landesregierung under grant PN 4067 “Unkonventionelle Forschung (UFO); Wiederlesbarmachung steirischer Römersteine mithilfe optischer Messtechniken”.