Daejeon, July 27–31
Digital methods in the humanities increasingly depend on the availability of structured and semantically linked data. Projects such as DataGEMShttps://datagems.eu/ aim to make data more accessible and reusable by allowing users to query complex data collections by establishing such links. These infrastructures enable connections between sources across different periods, languages, and formats, turning isolated corpora into linked data spaces that support new types of analysis. Encyclopedias provide a valuable case in this context: They capture the organization of knowledge at specific moments in history while revealing subtle knowledge shifts over time. Outliers in this context are inherently difficult to define: they emerge within a corpus that is deliberately designed to be coherent and standardized, yet is produced by multiple authors and shaped by differing editorial practices. Precisely because encyclopedias strive for a high degree of conceptual and stylistic uniformity, even small deviations become analytically meaningful, where they can reveal shifts in perspective or knowledge.
This contribution builds on these ideas. Using a semantically aligned corpus of nineteenth-century German encyclopedias employed in the context of DataGEMS, we apply embedding-based detection of conceptual outliers. Such deviations can point to changes in meaning or perspective and offer a data-driven way to study how descriptions evolve across editions and historical contexts.
This study builds on the tradition of conceptual history (Begriffsgeschichte), which views key terms as historically contingent and continuously redefined in response to social and political contexts (Koselleck 2006, Williams 1977). Humanities scholars view encyclopedias as cultural objects that not only reflect the values and power structures of their time but also help enforce them by deciding what counts as legitimate knowledge and how it should be presented (Yeo 2001, Loveland 2019).
In nineteenth-century Germany, the Konversationslexikon genre played a particularly influential role in the formation of the educated bourgeoisie (Bildungsbürgertum). Works such as Brockhaus and Meyer targeted the growing middle class by offering practical knowledge for everyday conversation while simultaneously teaching the implicit cultural codes necessary for social advancement (Spree 2014, Michel and Herren 2007). These encyclopedias functioned as instruments of sociable education, standardizing language, suppressing alternative meanings, and performing cultural authority for a broad readership. Confessional and thematic differences further differentiated the major encyclopedias: Brockhaus excelled in the humanities but offered weaker coverage of technical subjects, Meyer competed by strengthening technical and scientific content, and Herder’s Conversations-Lexikon served as the Catholic counterpart to the more Protestant tradition (Collison and Preece 2026).
As sources for historical semantics, encyclopedias fix the official meaning of words and concepts at one moment. This makes them invaluable for seeing how ideas are defined (and re-defined) over time:
“The diachronic comparison of lexicons [and encyclopedias] allows researchers to empirically demonstrate the repeatability of semantics and, at the same time, its innovation. [They] are therefore indispensable sources for any attempt to reconstruct the pace of change in the history of concepts” (translated from Koselleck 2006, p. 97).
In contrast to most work on semantic change in computational linguistics, which infers shifting meanings from usage patterns or distributional context (Wevers and Koolen 2020, Tahmasebi et al. 2021), this contribution examines explicit definitional variation across encyclopedic entries. Distributional methods excel at capturing gradual, usage-driven shifts from broad textual contexts. However, these methods are still confined to the specific corpus under study. As a result, detected changes may partly arise from corpus-specific variation (for example, shifts in genre balance, domain coverage, register, or sampling biases) rather than genuine semantic evolution. Even within that bounded corpus, the inferences remain approximations (Kutuzov et al. 2022).
Encyclopedias provide intentional, condensed statements of "what X is" at a given historical moment, authored by the editors with the explicit aim of defining and standardizing knowledge. Unlike inferred meanings from usage, definitional entries capture deliberate conceptualizations. Our embedding-based outlier detection approach therefore serves as a complementary, top-down lens for examining how knowledge is actively reformulated and explicitly defined in historical reference works; thereby contributing to both conceptual history and cultural analytics of encyclopedic discourse.
Several approaches to outlier detection exist, but many face challenges in high-dimensional, low-sample-size (HDLSS, Large p small n) settings. Distance-based methods flag points far from cluster centroids but can fail when distances concentrate in high dimensions. Density-based methods rely on local neighborhoods, which may be too sparse to estimate reliably when n is small. Projection methods reduce dimensionality before detecting anomalies, but they require sufficient structure to project onto, which may be unstable for very small clusters (Ahn 2019, Cavalheiro 2023, Nakayama 2024). We use an adapted distance-based approach, robust z-scores, which replace the mean and standard deviation with the median and MAD, motivated by Rousseeuw and Hubert (2018).
The corpus consists of 6 "conversational”, meaning general, encyclopedias from 1809 to 1911 Germany (cf. Table 1) plus a contemporary encyclopedia, i.e. the German Wikipedia. These encyclopedias cover all aspects of knowledge, including noteworthy people and places, but also objects and abstract concepts. They have been pre-processed within the frame of EncycNet,https://encycnet.github.io/ which included a cross-encyclopedic alignment of all the entries in the corpus (Hagen et al. 2022). In essence, using Wikipedia as a semantic baseline, EncycNet identified which entries belong together because they describe the same thing.
Our study aims to detect unusual conceptual entries in historical encyclopedias using embeddings. From EncycNet, we gather all sets of conceptually aligned entries, including its German Wikipedia equivalent.
We only look at concepts with 4 entries in at least 4 encyclopedias, where additionally, all entries must at least contain 100 tokens. This leaves us with 447 concept clusters. We then adopt a robust distance-based approach (Rousseeuw and Hubert, 2018). We first embed the text of each entry using a sentence-transformer model.https://huggingface.co/sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 For each concept, we compute a centroid of its embeddings across encyclopedias, which represents the typical conceptualization. For every entry, we calculate the cosine distance to its concept centroid:
where ei is the embedding of entry i and c is the centroid of all entries for that concept. To detect unusual entries, we normalize each entry’s distance using a robust z-score:
where D is the set of distances from entries to their centroid, and ε is a small constant for numerical stability. The median absolute deviation (MAD) is defined as:
An entry is flagged as an outlier if its robust z-score exceeds a predefined threshold. In this experiment, we set the threshold to 3.5 according to our distribution of scores to achieve a reasonable outlier percentage (Davies and Gather 1993), which also aligns with the recommendation of Iglewicz and Hoaglin (1993, p. 12).
By computing a centroid in embedding space and measuring cosine distances of each entry to this centroid, we obtain a simple yet informative representation of deviation from the typical conceptualization. Furthermore, by using a robust z-score based on median and MAD rather than mean and standard deviation, we reduce the influence of extreme values and achieve more stable estimates even in very small clusters.
The results are a list of concept clusters, where at least one of the entries is flagged as an outlier. We looked at 2238 entries split into 447 concepts. We found 146 outlier entries, which is about 6.5% of all entries in question.
Looking further into the specific entries marked as outliers, we can find different categories. First are entries that are unusual by design, which are, for example, supplement entries, entries from the appendix, or entries that for the most part consist of an image with a long image caption. Second, we also find errors from the original automatic alignment method, or entries, where the aligned Wikipedia article is not an exact fit. Third, we find outliers from encyclopedias which are unusual by design, such as the Damen Conversationslexikon (DamenLex). The final category of results are then entries, where there is no explanation other than a significantly different description given, indicative of a shift in knowledge. In the following, we will highlight some examples from the third and fourth categories. We add the full list of outliers (disregarding the false positives mentioned above) in the Appendix.
It comes to no surprise that many of the found outliers originate from the DamenLex, which is still a general “conversational” encyclopedia, but is tailored towards a female reader. This aligns with our previous quantitative analysis of the DamenLex, which identified gendered assumptions in its female biographical emphasis and stylistic features (Ketzan et al. 2022). The embedding-based outliers further reveal which entries deviate the most from the entirety of the DamenLex, or rather which entries differ from other descriptions way too much to appeal to women.
We find two patterns: one, the usage of metaphors or impersonizations of entire concepts, and two, literary passages used to create an impression of people or places, usually highly emotional. For the former pattern, for example, the concept of Stunde (hour) is in its entirety described as a mystical female entity, Nachtigall (nightingale) as a male singer, Liebe (love) is described as a plant, or Religion described as light. Concerning the latter pattern, we highlight most notably entries on William Shakespeare and Jean Jacques Rousseau, as well as Ulm and Thuringia. From the output it is notable that only entries on men tend to have this literary character.
Here, we find individual outliers that could be one of the historical encyclopedias, where either we find a special focus of the concept or entirely new knowledge in an outlier, or, the outlier is actually the Wikipedia entry, indicative of a general shift in semantics, tone, or knowledge. One example would be Antoinette Bourignon. It is stated in the 1809 entry that she fled to a desert for instance (she was born in Flandern) or that she was given a place to settle by the Archbishop of Cambrai, which none of the other entries mention. Similarly, Gottfried August Bürger in his 1809 entry is praised as “the favourite poet of the nation.” Generally speaking, there are noticeable differences in people’s biographies found in this category, which is sometimes simply the case because of the person still being alive at the time of an entry being written. Concepts can also be outliers: the 1905 entry on Knall (bang) for example. Only that entry is entirely focused on very technical details pertaining to firearms.
We may also look at Wikipedia outliers; here we find, for example, Zigeuner (gypsy), and also noteworthy personalities such as Elisabeth Petrowna for instance, who was painted as a combination of moody, salacious, or passive and insignificant in historical entries. Wikipedia on the other hand also attributes her passiveness to sickness, which none of the historical entries mention. Some outliers are also distinct senses of a word seemingly lost today, for example Bekleidung 1905 (clothing) in a strictly military sense, or Einsalzen 1854 (salting) used just for animal food.
To better understand the added value of our definitional outlier detection approach, we conducted a complementary experiment using word embeddings following the exact approach of Hamilton et al. (2016). We trained separate embedding models on (i) all historical encyclopedias combined versus German Wikipedia to capture diachronic change, and (ii) all works versus the Damen Conversations Lexikon alone to capture synchronic stylistic variation. After performing orthogonal Procrustes alignment, we examined both cosine distances between all terms (to see where our already flagged terms rank) and investigated nearest neighbors.
The two methods overlap on several outliers, including Stunde (hour) and Zigeuner (gypsy). However, the underlying reasons for term deviations may differ. In the DamenLex, Stunde appears in a concrete, spatial sense (e.g., as a measure of distance between cities), whereas our entry-based outlier detection primarily highlights its metaphorical and personified usage. For Zigeuner, the historical nearest neighbors point to the description of the people, while the Wikipedia nearest neighbors point to the description of the negative connotation—aligning with the entry outlier approach. At the same time, several outliers identified by our definitional approach do not appear as divergent in the contextual embedding analysis.
A clear practical advantage of the entry-alignment approach emerges with named entities and multi-word expressions. The Hamilton-style method struggles with historical spelling variation (e.g., Jan/Johann/Johannes Hus/Huß), often requiring prior entity detection and linking to track specific concepts reliably. In contrast, the semantic alignment provided by EncycNet operates at the level of complete entries, where a single alignment suffices despite orthographic differences.
Several limitations should be considered when interpreting the results. First, the size of many conceptual clusters is exceptionally small, even for HDLSS. This low sample size limits the statistical stability of any distance- or embedding-based outlier detection and means that flagged outliers should be interpreted as candidates for further qualitative analysis rather than definitive anomalies. Second, our method assumes that most entries represent a “typical” conceptualization. In cases where the underlying entries are themselves highly heterogeneous, this assumption may not hold. Third, some outliers also arise from stylistic variation or alignment errors. In this study, we addressed this by manually categorizing flagged outliers into groups (e.g., design-related anomalies, alignment errors, and potential conceptual shifts) and qualitatively interpreting the latter as candidates for deeper analysis. This manual disentanglement step is of course labor-intensive and relies on domain expertise, however computationally disentangling style and format effects from genuine conceptual drift remains highly challenging and constitutes an important direction for future work. That same difficulty also still applies to semantic shift detection in usage-based approaches, where distinguishing other corpus effects from true conceptual change often requires manual inspection (Kutuzov 2022). Our method therefore does not aim to replace any qualitative or usage-based work but serves as a complementary method that highlights promising candidates for closer scrutiny.
The experiment illustrates the potential of semantically linked data for large-scale cultural analyses, which is one of the core goals of DataGEMS, where a multitude of sources will be aligned over the course of its duration. Future work could improve preprocessing to reduce false positives, adapt outlier thresholds, expand the dataset to a multilingual setting, and explore complementary embedding models or temporal/contextual representations to better capture subtle shifts. Beyond comparative settings, the approach could also be applied within a single knowledge community that explicitly strives for coherence, such as Wikipedia, where detecting internal outliers may help identify deviations in tone as well as different approaches to conceptualization.
This work has been fully supported by DataGEMS, funded by the European Union's Horizon Europe Research and Innovation programme, under grant agreement No 101188416.