DH 2026

Daejeon, July 27–31

Fri, July 3114:00–15:30S009101-102
Long Paper

Silenced in the Data: Epistemic Injustice as Challenge for the AI-Assisted Analysis of Abuse in the Catholic Church

Thomas Schmidt
Media Informatics Group, University of Regensburg, Germany · thomas.schmidt@ur.de
Magdalena Hürten
Chair of Pastoral Theology and Homiletics, University of Regensburg, Germany · magdalena.huerten@theologie.uni-regensburg.de
Ute Leimgruber
Chair of Pastoral Theology and Homiletics, University of Regensburg, Germany · ute.leimgruber@theologie.uni-regensburg.de
Christian Wolff
Media Informatics Group, University of Regensburg, Germany · christian.wolff@ur.de

Introduction

There is a long tradition of theological research within Digital Humanities (DH), reaching back to Roberto Busa’s Index Thomisticus, one of the foundational projects of the field (Jones 2018). More recently, this strand of research has been consolidated under umbrella terms such as computational theology and digital theology (Nunn/Van Oorschot 2024). Drawing on a literature review and the references compiled in the bibliography, we present results from an ongoing DH-based reappraisal of the sexual abuse in the Catholic Church. While public debate and scholarly research have largely focused on the abuse of children, the experiences of adult women and adolescent girls have been systematically marginalized (Hübel 2020; Gott 2022; Leimgruber 2026). This neglect can be understood as a form of epistemic injustice (Fricker 2007; Hürten 2025). Using AI-supported methods, we present results of a project investigating the German discourse on the sexual abuse of women in media coverage and official Catholic documentation.

The project “Von Epistemic Injustice zu Epistemic Awareness. Prozessbasierte Methodenforschung zu Missbrauch an erwachsenen Frauen in der katholischen Kirche” (From Epistemic Injustice to Epistemic Awareness. Process‑based Methodological Research on Abuse of Adult Women in the Catholic Church) is funded by the DFG (German Research Association, project number 539293243): https://missbrauchsmuster.de/forschen/projekte/forschungsprojekt-von-epistemic-injustice-zu-epistemic-awareness-prozessbasierte-methodenforschung-zu-missbrauch-an-erwachsenen-frauen-in-der-katholischen-kirche/

Particular attention is paid to what we term "hidden patterns", recurring but largely invisible structures in language, behavior, and institutional practices that systematically obscure certain forms of violence and render specific victim groups imperceptible in documentation and discourse (Leimgruber 2026, 88–91). Our goal is to develop corpora and computational approaches that make these marginalized narratives analytically accessible. The contributions of this paper are threefold: (1) the creation of an annotated corpus of five official abuse reports comprising 1,658 expert annotations and an associated annotation scheme; (2) an evaluation of zero-shot relevance detection using generative large language models (LLMs); and (3) results from fine-grained annotation classification using discriminative LLMs. Beyond their methodological implications, these results directly engage with the conference theme of Beyond Patterns: Interpretation with Small Data. As we demonstrate, the marginalization of women within abuse discourse produces severe data sparsity and class imbalance, posing distinctive challenges for annotation practices and machine learning alike.

Corpus and Relevance Annotation

Alongside media coverage, social media content, survivor

The term survivor is widely preferred in international abuse research because it avoids positioning individuals in a passive or objectified role—a connotation that the term victim can inadvertently reinforce.

testimonies, and scholarly literature, abuse reports constitute a key text type for analyzing the discourse on sexual abuse. These reports have been commissioned by dioceses or bishops’ conferences and investigate sexual abuse committed by members of the Church after the Second World War and up to the present. In Germany, more than 30 official abuse reports have been published and made publicly available online. From this body of material, we annotated a representative sub-corpus of five reports commissioned by the dioceses of Bozen–Brixen (Wastl et al. 2025), Essen (Dill et al. 2023), Hildesheim (Hackenschmied and Mosser 2017), Cologne (Gercke et al. 2021), and the Swiss bishops’ conference (Bignasca et al. 2023). The selection of reports was based on representative criteria designed to capture the diversity of existing investigations: one of the earliest published reports and one of the most recent; reports with a primarily legal orientation and those grounded in social‑science approaches; reports focusing on the abuse of minors by clerics as well as those also considering abuse of adults and abuse committed by church employees in general. Table 1 summarizes the descriptive statistics of these reports.

Table 1. General corpus statistics of five abuse reports.

The entire corpus consists of 2,362 PDF pages and almost 730,000 tokens (after removing noise) with the longest report having 915 pages. We aim to develop AI systems to first identify relevant sections. Relevance is determined as sections that deal with discourse and reporting of the sexual abuse of specifically adult women and/or girls above the age of 14. In these sections, we want to perform further fine-grained classification according to a complex annotation scheme. An expert annotator manually identified all relevant sections in the PDFs, resulting in approximately 6% of pages being marked as relevant, reduced to just the relevant content making 3.6% of the overall text. These annotations indicate that the abuse of adult women and adolescent girls is rarely discussed in the reports.

Relevance Detection

We employ zero-shot relevance detection using an instruction-tuned LLM (Gemma 3:12b) (Gemma Team et al. 2025). Zero-shot detection refers to the LLM method in which a model assigns classes (here: relevance) to a text unit without having been explicitly trained on labeled examples. This approach was chosen due to the large class imbalance and the limited availability of labeled data. We prioritize high-recall detection because hidden patterns in the texts make it likely that relevant passages are missed even in expert annotations, and relying solely on pre-labeled training data could risk reproducing the very epistemic injustice under investigation. Zero-shot classification has proven particularly effective for rare, implicitly articulated, and semantically complex categories, as it leverages contextual meaning rather than surface-level lexical features (Zhang et al. 2019; Schopf et al. 2023). Our pipeline follows a human-in-the-loop approach: expert-labeled positive examples are analyzed to better understand hidden patterns, identify missed relevant sections, and evaluate the limitations and benefits of generative models. For evaluation, all texts were split into chunks of 10,000 characters, the maximum size of all relevant sections, and chunks containing relevant content were labeled accordingly (Table 1). Through experimentation, we found that a simplified and strict prompt, marking passages as relevant only when certain, produced the best results. Figure 1 presents an English translation of the original German prompt. Two additional prompts were used to collect the model’s explanations and reasoning for further in-depth analysis. Tables 2 and 3 summarize the key evaluation metrics.

Figure 1. Relevance detection prompt (translated from German).

Table 2. Relevance prediction cross table for the LLM and the gold standard (GS) annotation on all chunks.

Table 3. Evaluation metrics overall (accuracy, F1) and specifically for the positive “relevant” class (precision, recall).

The evaluation results indicate a tendency for the model to overclassify chunks as relevant, resulting in low precision but higher recall for the positive minority class. Analysis of false positives and the model’s reasoning yielded several insights:

  • General prompt pressure led to overclassification in unclear and vague cases though we could lower this in the process through a stricter, binary prompt (figure 1).
  • The model often labeled any general mentioning of women as relevant, particularly when adult women were reporting abuse experienced as children.
  • Chunks that were too long frequently mixed signals, making classification more difficult.

Overall, the vague and unclear language of the reports about age and situation of survivors, itself a manifestation of hidden patterns, poses challenges to the model. However, we did indeed find relevant texts in the false positives showing possible benefits. We intend to use the model as a high-recall filter rather than a definitive classifier, prioritizing sensitivity over precision to minimize missed cases and learn more about the discourse.

Fine-grained Annotation and Classification

After the identification of relevant passages, we intend to perform more detailed fine-grained analysis of the texts. For this purpose, we developed a complex hierarchical annotation scheme in an iterative process (Reiter 2020). The scheme consists of 16 main categories and 122 sub-categories covering all relevant research aspects of the topic, including information about survivors and accused persons, the actions of Church authorities, and the framing of the crime (see figure 2 for an example annotation). We annotated all sections marked as relevant by the expert annotator according to the scheme resulting in 1,658 annotations of varied lengths. For annotation, we used the CATMA tool (Gius et al. 2025) which is an established tool used in DH research for the flexible annotation of large form texts and shows wide usage in similar projects (Dennerlein et al. 2023; Heyns and van Zaanen, 2024; Guhr et al. 2025).

Figure 2. Annotation example.

Figure 2 provides insight into the annotation practice (Bignasca et al. 2023, 81). Annotated elements include: in blue the accused Hansjörg Vogel (gender: male, status: cleric), in light green the victim (gender: female, age: adult), in brown the description of the crimes (here: consequences of pregnancy, annotated as description of the crimes_reproductive), in yellow sexual acts including penetration (must have happened previously), in dark green the bishop’s reaction in the form of his resignation, and in pink the year of the reaction. Please note that the original annotation was done for German text and this is an English translation.

After narrowing the texts to relevant passages, class imbalance and the lack of labeled data are less severe for individual annotation classes. Accordingly, we treat discriminative transformer-based language models as state-of-the-art for classification tasks. Nevertheless, an imbalance persists, as most of the text within relevant passages does not correspond to any annotation category. We formulate the task as binary classification between a given category and non-annotated text, splitting both annotations and relevant text into sentences and assigning the corresponding labels (resulting in 1,422 non-annotated sentences). For evaluation, we tested the largest German-language BERT models, including gbert-large (Chan et al. 2020) and ModernGBERT (Wunderle et al. 2025). To address class imbalance, we experimented with strategies such as class weighting, oversampling, back-translation, and paraphrasing (Taskiran et al. 2025). Without such optimization, the models consistently collapsed toward the negative majority class, in contrast to the generative LLMs discussed in section 3. Table 4 presents the results of the best-performing models in a 5×5 evaluation for the five most frequently annotated main categories.

Table 4. Classification results.

PR-AUC values above 0.7 indicate good integration of the minority class, and balanced accuracy values of around 0.8 are encouraging. Our evaluations showed various insights and best practices for small data settings. The best configuration with the model bert-base-german-cased

https://huggingface.co/google-bert/bert-base-german-cased

includes the application of class weights and oversampling through back-translation (German à English à German).

Further model information: Fine-tuning for 4 epochs, learning rate of 4e-5, Adam optimizer, trained on a A100 GPU. Back-translation was performed with sentencepiece (https://github.com/google/sentencepiece)

No oversampling or only oversampling with direct duplicates pushed the model to learn classes by memorizing frequent words which could be best avoided by including artificial variance via back-translation while paraphrasing complicated source text. More information about the fine-grained annotation and classification can be found in Hürten et al. (2026).

Conclusion

This study marks important components of a larger effort to uncover the marginalized voices in German discourse on sexual abuse in the Catholic Church. All annotations, data and results are publicly available on a GitHub repository.

https://github.com/lauchblatt/Epi_Epa

We intend to develop and optimize pipelines like the one presented for other text types such as media texts and survivor testimonies. With the limited set of texts, we have already gained insights through annotation scheme development and annotation analysis. Our machine learning results are centered around the challenge of small data sets and class imbalance. The high-recall relevance detection using generative LLMs is encouraging, and we intend to explore the models further prioritizing human judgement before AI in a human-in-the-loop approach, with additional annotations as well as advanced few-shot-learning (giving the model a small set of labeled examples to improve classification). The fine-grained classification model shows promising results, and we look forward to applying optimized models on large unlabeled data to further our advances in turning “silenced data” into analytically and socially meaningful knowledge.

Sexual abuse has been systematically covered up in the Catholic Church, and the abuse of women continues to be formally recognized only under narrow conditions. Uncovering “hidden knowledge” about such cases therefore is crucial to gain insight into the phenomenon, develop preventive measures and promote (epistemic) justice for the survivors. To our knowledge, there is currently little comparable computational work focused specifically on systematically identifying and analyzing marginalized survivor perspectives within the Catholic Church abuse discourse, highlighting the methodological and exploratory contribution of this project to both DH and computational social analysis. By demonstrating how computational methods can surface epistemic injustice embedded in institutional documentation, this work contributes to broader DH discussions about the ethical dimensions of data interpretation and the role of digital methods in amplifying marginalized voices.

References
  1. Bignasca, Vanessa, Lucas Federer, Magda Kaspar and Lorraine Odier. 2023. “Bericht zum Pilotprojekt zur Geschichte sexuellen Missbrauchs im Umfeld der römisch-katholischen Kirche in der Schweiz seit Mitte des 20. Jahrhunderts. Bern. https://doi.org/10.5281/zenodo.8315772
  2. Chan, Branden, Stefan Schweter and Timo Möller (2020) “German`s Next Language Model“, in D. Scott, N. Bel, und C. Zong (Edt.) Proceedings of the 28th International Conference on Computational Linguistics. COLING 2020, Barcelona, Spain (Online): International Committee on Computational Linguistics, pp. 6788–6796. https://doi.org/10.18653/v1/2020.coling-main.598
  3. Dennerlein, Katrin, Thomas Schmidt and Christian Wolff. 2023. “Computational Emotion Classification for Genre Corpora of German Tragedies and Comedies from 17th to Early 19th Century.” Digital Scholarship in the Humanities 38(4): 1466–1481.
  4. Dill, Helga, Malte Täubrich, Peter Caspari, Tinka Schubert, Gerhard Hackenschmied, Elan Pinar and Elisabeth Helming. 2023. “Aufarbeitung sexualisierter Gewalt im Bistum Essen: Fallbezogene und gemeindeorientierte Analysen.“ www.bistum-essen.de/fileadmin/relaunch/Bilder/Soziales_und_Hilfe/Sexueller_Missbrauch/ipp/IPP_Studie_Bistum_Essen.pdf (accessed: May 1, 2025).
  5. Fricker, Miranda. 2007. Epistemic Injustice: Power & the Ethics of Knowing. Oxford: Oxford University Press.
  6. Gemma Team et al. 2025. “Gemma 3 Technical Report.” https://arxiv.org/abs/2503.1978
  7. Gercke, Björn, Kerstin Stirner, Corinna Reckmann and Max Nosthoff-Horstmann. 2021. “Pflichtverletzungen von Diözesanverantwortlichen des Erzbistums Köln im Umgang mit Fällen sexuellen Missbrauchs von Minderjährigen und Schutzbefohlenen durch Kleriker oder sonstige pastorale Mitarbeitende des Erzbistums Köln im Zeitraum von 1975 bis 2018: Verantwortlichkeiten, Ursachen und Handlungsempfehlungen.“ mam.erzbistum-koeln.de/m/2fce82a0f87ee070/original/Gutachten-Pflichtverletzungen-von-Diozesanverantwortlichen-im-Erzbistum-Koln-im-Umgang-mit-Fallen-sexuellen-Missbrauchs-zwischen-1975-und-2018.pdf (accessed: May 1, 2025).
  8. Gius, Evelyn, Jan Christoph Meister, Malte Meister, Marco Petris, Dominik Gerstorfer, Mari Akazawa and Stephanie Messner. 2025. “CATMA (7.2.0)“. Zenodo. https://doi.org/10.5281/zenodo.1470118
  9. Gott, Chloë K. 2022. Experience, Identity & Epistemic Injustice within Ireland’s Magdalene Laundries. Bloomsbury studies in religion, gender, and sexuality. London: Bloomsbury.
  10. Guhr, Svenja, Huijun Mao and Fengyi Lin. 2025. “Rethinking Scene Segmentation. Advancing Automated Detection of Scene Changes in Literary Texts”. In Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2025), edited by Anna Kazantseva, Stan Szpakowicz, Stefania Degaetano-Ortlieb, Yuri Bizzoni, und Janis Pagel. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.latechclfl-1.8
  11. Hackenschmied, Gerhard and Peter Mosser. 2017. “Untersuchung von Fällen sexualisierter Gewalt im Verantwortungsbereich des Bistums Hildesheim: Fallverläufe, Verantwortlichkeiten, Empfehlungen.“ www.praevention.bistum-hildesheim.de/fileadmin/dateien/PDFs/Pressetexte/IPP_Muenchen_Gutachten_Bistum_Hildesheim.pdf (accessed: May 1, 2025).
  12. Heyns, Nuette and Menno van Zaanen. 2025. “Annotated Mystery Narratives Data Set”. Journal of Open Humanities Data 11(1). https://doi.org/10.5334/johd.391
  13. Hübel, Sylvia. 2020. “Herstory of Epistemic Injustice: Wo/men’s Silencing in the Catholic Church”. Journal of the European Society of Women in Theological Research 28: 127–152. https://doi.org/10.2143/ESWTR.28.0.3288485.
  14. Hürten, Magdalena. 2025. Dem Schweigen zuhören. Die Bedeutung des Konzepts der epistemic injustice für die Forschung zu Missbrauch an erwachsenen Frauen in der katholischen Kirche: Fallstudie zur Gründungsgeschichte der St. Franziskusschwestern Vierzehnheiligen. Religion - Geschlecht - Körper | religion - gender - bodies 2. Baden-Baden: Karl Alber. https://doi.org/10.5771/9783495992302
  15. Hürten, Magdalena, Thomas Schmidt, Ute Leimgruber and Christian Wolff. 2026. “Sexual Abuse in the Catholic Church: Annotation and Machine Learning of Hidden Patterns in Abuse Reports.” Digital Humanities Benelux 2026 (DHBenelux 2026). Maastricht, Netherlands. https://doi.org/10.5281/zenodo.19663945
  16. Jones, Steven E. 2018. Roberto Busa, S. J., and the Emergence of Humanities Computing: The Priest and the Punched Cards. Routledge.
  17. Leimgruber, Ute. 2020. “‘Hidden Patterns‘ – Überlegungen zu einer machtsensiblen Pastoraltheologie.“ ET-Studies 11: 207–224.
  18. Leimgruber, Ute. 2026. Missbrauchsmuster: Was Missbrauch an Frauen ermöglicht und warum er nicht als solcher erkannt wird. Ostfildern: Matthias Grünewald.
  19. Leimgruber, Ute and Doris Reisinger. 2021. “Sexueller Missbrauch oder sexualisierte Gewalt?“ feinschwarz.net. www.feinschwarz.net/sexueller-missbrauch-oder-sexualisierte-gewalt-ein-einspruch (accessed: June 12, 2025).
  20. Nunn, Christopher A. and van Oorschot, Frederike (Edt.). 2024. Kompendium Computational Theology, Bd. 1: Forschungspraktiken in den Digital Humanities. Heidelberg: heiBOOKS. https://doi.org/10.11588/heibooks.1459
  21. Reiter, Nils. 2020. “Anleitung zur Erstellung von Annotationsrichtlinien.“ Reflektierte algorithmische Textanalyse, edited by Nils Reiter, Axel Pichler and Jonas Kuhn, 193-202. Berlin/Boston: De Gruyter. https://doi.org/10.1515/9783110693973
  22. Taskiran, Salimkan Fatma et al. 2025. "A comprehensive evaluation of oversampling techniques for enhancing text classification performance.” Scientific Reports, 15(1), 21631. https://doi.org/10.1038/s41598-025-05791-7
  23. Schopf, Tim, Daniel Braun and Florian Matthes. 2023. “Evaluating Unsupervised Text Classification: Zero-shot and Similarity-based Approaches.” Proceedings of the 2022 6th International Conference on Natural Language Processing and Information Retrieval (NLPIR '22). Association for Computing Machinery, New York, NY, USA, 6–15. https://doi.org/10.1145/3582768.3582795
  24. Wastl, Ulrich, Martin Pusch, Nata Gladstein and Philipp Schenke. 2025. “Sexueller Missbrauch Minderjähriger und erwachsener Schutzbefohlener durch Kleriker im Bereich der Diözese Bozen-Brixen von 1964 bis 2023: Verantwortlichkeiten, systemische Ursachen und Empfehlungen.“ westpfahl-spilker.de/wp-content/uploads/2025/01/Gutachten_Dioezese-Bozen-Brixen_DE-1.pdf (accessed: May 1, 2025).
  25. Wunderle, Julia et al. 2025. “New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models.” https://arxiv.org/abs/2505.13136
  26. Zhang, Jingqing, Piyawat Lertvittayakumjorn and Yike Guo. 2019. “Integrating Semantic Knowledge to Tackle Zero-shot Text Classification.” Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 1031–1040, Minneapolis, Minnesota. Association for Computational Linguistics.