DH 2026

Daejeon, July 27–31

Poster

Epistemic Injustice and Ecclesial Discourse: A Digital Humanities Perspective on the Sexual Abuse of Women in the Catholic Church

Magdalena Hürten
Chair of Pastoral Theology and Homiletics, University of Regensburg, Germany · magdalena.huerten@theologie.uni-regensburg.de
Thomas Schmidt
Media Informatics Group, University of Regensburg, Germany · thomas.schmidt@ur.de
Ute Leimgruber
Chair of Pastoral Theology and Homiletics, University of Regensburg, Germany · ute.leimgruber@theologie.uni-regensburg.de
Christian Wolff
Media Informatics Group, University of Regensburg, Germany · christian.wolff@ur.de

Introduction

In this poster contribution, we present an overview of an ongoing Digital Humanities (DH) research project

The project “Von Epistemic Injustice zu Epistemic Awareness. Prozessbasierte Methodenforschung zu Missbrauch an erwachsenen Frauen in der katholischen Kirche” (From Epistemic Injustice to Epistemic Awareness. Process‑based Methodological Research on Abuse of Adult Women in the Catholic Church) is funded by the DFG (German Research Association, project number 539293243): https://missbrauchsmuster.de/forschen/projekte/forschungsprojekt-von-epistemic-injustice-zu-epistemic-awareness-prozessbasierte-methodenforschung-zu-missbrauch-an-erwachsenen-frauen-in-der-katholischen-kirche/

situated within the long tradition of theological scholarship in DH, tracing back to Roberto Busa’s pioneering collaborations with IBM (Jones 2018). The project aims to advance the understanding of discourses surrounding the sexual abuse of adult women and adolescent girls (aged 14 and above) in the German speaking Catholic Church, with a particular focus on identifying “hidden patterns” of silencing and neglect (Leimgruber 2026, 47) that constitute forms of epistemic injustice (Fricker 2007; Hürten 2025). Following Fricker (2007, 1), we understand epistemic injustice as “consisting, most fundamentally, in a wrong done to someone specifically in their capacity as a knower.” The project integrates multiple methodological components and research objectives and is in line with similar current research exploring the application of Large Language Models (LLMs) in DH (e.g. Dennerlein et al. 2023; Hellwig et al. 2024):

  • The collection and preparation of text corpora relevant for the topic of sexual abuse of women in the catholic church e.g. media reports, survivor testimonies, abuse reports.
  • The development of annotation schemes and the annotation of sub-corpora.
  • AI-assisted relevance detection with LLMs to identify the (rare) passages in documents dealing with the sexual abuse of adult women/adolescent girls.
  • Fine-grained classification with LLMs according to the annotation scheme.
  • Large-scale analysis of documents in a reflective and critical human-in-the-loop process.

We present selected results from each component to demonstrate the project’s potential and to outline its current direction through illustrative research findings.

Abuse Reports Analysis

In a first pilot study with five abuse reports (2,363 PDF pages) commissioned by different dioceses and bishops conferences in German speaking regions and published between 2017 and 2025,

Most of the German dioceses have by now commissioned independent researchers or law firms to investigate abuse cases within the diocese. The research design and scope of the reports vary greatly. Yet, the majority of reports focusses on the abuse of minors by clerics examining the period from after the Second World war to the present and drawing primarily on archival material provided by the dioceses and, in some cases, supplemented by interviews with survivors, church officials and other contemporaries. The reports are generally made available to the public online as is the case with the five reports that have been analyzed for the pilot study. All data were anonymized by the editors prior to the publication in accordance with European data protection regulations.

we marked relevant texts (passages dealing with adult women/adolescent girls), developed a fine-grained annotation scheme and performed annotations with the CATMA tool (Gius et al. 2025) resulting in 1,658 annotations regarding survivor and accused information, the framings used to address abuse, described nature of the abuse, or actions of involved actors (divided in survivors

The term survivor is widely preferred in international abuse research because it avoids positioning individuals in a passive or objectified role—a connotation that the term victim can inadvertently reinforce.

, accused, Church officials and social network of the survivors).

Analysis of the annotations has produced initial insights into the silencing of women’s experiences in the reports. It reveals that even reports focused on the abuse of minors document abuse against adults. However, these cases scarcely register in the public reception of the reports (see also Leimgruber 2026, 96f.). Moreover, the case-based analysis shows that several perpetrators abused victims from more than one age group. Because each case is defined by a single perpetrator, individual cases may include multiple victims of different ages. In seven of the cases that were annotated as relevant, the ages of all victims could be determined, and each of these involved both minor and adult victims. Two additional cases likewise included both minors and adults but also involved further victims whose ages could not be established (Figure 1). This overlap of survivor age groups refutes the assumption of predominantly pedocriminal offenders within the Catholic Church

This assumption was already refuted by Dreßing et al. (2018, 128), who demonstrated that pedophilic preference disorders could only be assumed for a subset of offenders within the Catholic Church.

and questions the marginalization of adult women survivors in the abuse discourse.

Figure 1. Age group combinations per relevant case

Only 3.4% of the original texts were manually marked as relevant and annotated (highlighting the neglect of this survivor group in the reports). This massive class imbalance and lack of labeled data pose a challenge to all machine learning tasks. In a first step, we aim to design a high-recall relevance detection system using generative AI (Zhang et al. 2019; Schopf et al. 2023). In experiments with an instruction-tuned LLM (Gemma 3:12b) (Gemma Team et al. 2025) we achieve macro F1-scores of 0.6 with high recall and low precision for the positive class.

Due to the sensitive nature of the material concerning sexual abuse in the Catholic Church, all LLM-based analyses were conducted on local university server infrastructure.

We use a human-in-the-loop process

This work is carried out either by experts with long-standing experience in researching abuse in the Catholic Church or by trained student assistants working in close collaboration with them. New team members receive guidance in self‑care practices, and emotionally demanding material is regularly addressed in collegial debriefings.

to screen the false positives and analyze suggestions and potential hidden patterns of the AI system. The separate task of fine-grained classification of relevant passages yields success using oversampling techniques and BERT models (Chan et al. 2020) (F1-scores ranging from 0.8 to 0.9).

Further model information: German BERT model with fine-tuning for 4 epochs, learning rate of 4e-5, Adam optimizer, trained on a A100 GPU. Optimization via class weights and oversampling through back-translation.

Web Analysis

Another important text type is public media reports on the web. Using the Google Search API

https://developers.google.com/custom-search/v1/overview?hl=de

with fitting terms, we gathered a corpus of over 10,000 HTML pages of different websites of the Catholic Church, major German media outlets and other specific pages dedicated to the topic. We developed a web-based annotation tool (figure 3)

The tool is based on Open Search and PostgresSQL.

and are currently in the process of annotating the relevance of documents to eventually implement a similar pipeline as for the abuse reports.

Figure 3. Screenshot of annotation tool. Annotators can perform advanced searches over scraped corpora and annotate relevance information.

To illustrate the potential of web-based research, we present results from a case study of Gottes-Suche.de

https://www.gottes-suche.de/

(“The search for God”), a website that manually aggregates reports on sexual abuse in the Catholic Church. We scraped all short reports available on the platform (approximately 10,000 texts) and conducted a semantic analysis. Figure 4 shows that references to women (and semantically related terms) increase over time, with a notable peak in 2019, indicating a growing awareness of this survivor group within media discourse.

Figure 4: Percentage of the usage of words like women throughout the collected media reports on Gottes-Suche.de from 2010 to 2024.

These example results from Gottes-Suche.de demonstrate the potential of large-scale analysis once sufficiently robust AI systems are in place. At the same time, given the ethical sensitivity of the research domain, it is essential to prioritize human expert judgment. We therefore continue to pursue a human-in-the-loop approach in which AI systems are iteratively improved through continuous reflection on and validation of their outputs. All data, annotations, analyses and more information about the AI application are made publicly available via a GitHub repository to support transparency and further research in this field.

https://github.com/lauchblatt/Epi_Epa

References
  1. Chan, Branden, Stefan Schweter and Timo Möller (2020) “German`s Next Language Model“, in D. Scott, N. Bel, und C. Zong (Edt.) Proceedings of the 28th International Conference on Computational Linguistics. COLING 2020, Barcelona, Spain (Online): International Committee on Computational Linguistics, pp. 6788–6796. https://doi.org/10.18653/v1/2020.coling-main.598
  2. Dennerlein, Katrin, Thomas Schmidt and Christian Wolff. 2023. “Computational Emotion Classification for Genrecorpora of German Tragedies and Comedies from 17th to Early 19th Century.” Digital Scholarship in the Humanities 38(4): 1466–1481. 
  3. Dreßing, Harald, Hans J. Salize, Dieter Dölling, Dieter Hermann, Andreas Kruse, Eric Schmitt and Britta Bannenberg. 2018. „Sexueller Missbrauch an Minderjährigen durch katholische Priester, Diakone und männliche Ordensangehörige im Bereich der Deutschen Bischofskonferenz (MHG-Studie).“ www.zi-mannheim.de/fileadmin/user_upload/downloads/forschung/forschungsverbuende/MHG-Studie-gesamt.pdf (accessed: July 22, 2025). 
  4. Fricker, Miranda. 2007. Epistemic Injustice: Power & the Ethics of Knowing. Oxford: Oxford University Press.
  5. Gemma Team et al. 2025. “Gemma 3 Technical Report.” https://arxiv.org/abs/2503.1978.
  6. Gius, Evelyn, Jan Christoph Meister, Malte Meister, Marco Petris, Dominik Gerstorfer, Mari Akazawa and Stephanie Messner. 2025. “CATMA (7.2.0)“. Zenodo. https://doi.org/10.5281/zenodo.1470118.
  7. Hellwig, Nils C., Jakob Fehle, Markus Bink, Thomas Schmidt and Christian Wolff. 2024. “Exploring Twitter Discourse with BERTopic: Topic Modeling of Tweets Related to the Major German Parties during the 2021 German Federal Election.” International Journal of Speech Technology 27(4): 901–921.
  8. Hürten, Magdalena. 2025. Dem Schweigen zuhören: Die Bedeutung des Konzepts der epistemic injustice für die Forschung zu Missbrauch an erwachsenen Frauen in der katholischen Kirche: Fallstudie zur Gründungsgeschichte der St. Franziskusschwestern Vierzehnheiligen (Religion - Geschlecht - Körper | religion - gender - bodies 2). Baden-Baden: Karl Alber. https://doi.org/10.5771/9783495992302
  9. Jones, Steven E. 2018. Roberto Busa, S. J., and the Emergence of Humanities Computing: The Priest and the Punched Cards. Routledge, 2018.
  10. Leimgruber, Ute. 2026. Missbrauchsmuster: Was Missbrauch an Frauen ermöglicht und warum er nicht als solcher erkannt wird. Ostfildern: Matthias Grünewald.
  11. Schopf, Tim, Daniel Braun and Florian Matthes. 2023. “Evaluating Unsupervised Text Classification: Zero-shot and Similarity-based Approaches.” Proceedings of the 2022 6th International Conference on Natural Language Processing and Information Retrieval (NLPIR '22). Association for Computing Machinery, New York, NY, USA, 6–15. https://doi.org/10.1145/3582768.3582795
  12. Zhang, Jingqing, Piyawat Lertvittayakumjorn and Yike Guo. 2019. “Integrating Semantic Knowledge to Tackle Zero-shot Text Classification.” Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 1031–1040, Minneapolis, Minnesota. Association for Computational Linguistics.