DH 2026

Daejeon, July 27–31

Workshop

Bringing Humanist Practices into AI – establishing and maintaining sovereignty when working with LLMs

Noah Jefferson Kim-Baumann
Humboldt University of Berlin, Germany · noah.jefferson.baumann.1@hu-berlin.de
Torsten Hiltmann
Humboldt University of Berlin, Germany · torsten.hiltmann@hu-berlin.de

Introduction & Context

The integration of Large Language Models (LLMs) into the digital humanities offers powerful capabilities for a wide spectrum of workflows. This workshop addresses the conference theme of “Engagement” by fostering a critical, active interaction between humanities scholars and AI. We believe that today, when AI agents can generate literature reviews and source corpora with a single click, it is more important than ever that digital humanists understand underlying technologies and display AI literacy. It is then scholar’s responsibility to reflect on how to maintain our own discipline practices when working with AI. Retrieval-Augmented Generation (RAG) explicitly exposes the possibilities and challenges of LLMs, making it the ideal laboratory for developing critical AI literacy. Using a custom RAG system adapted for the United Nations General Debate Corpus and digital humanities practices, this workshop moves beyond “chatting with data.” We aim to teach scholars to work with RAG systems as reflective research aids, ensuring that the “memory of the world” is not just retrieved, but critically interpreted. Participants will leave empowered to maintain their scholarly sovereignty by understanding and critically directing AI systems, rather than being directed by them. This workshop is supported by the NFDI4Memory Methods Innovation Lab.

Current State of Knowledge

RAG combines pre-trained parametric memory with non-parametric external knowledge, reducing hallucinations and improving factual consistency (Lewis et al. 2020). While RAG has rapidly evolved to include modular architectures and advanced retrieval strategies (Gao et al. 2024), its application in the humanities often lacks specific epistemological and disciplinary grounding. Recent surveys highlight that while RAG enhances reliability, it introduces new risks regarding explainability and bias that require trustworthy frameworks (Ni et al. 2025).

In historically-focused humanities research, the transition from traditional source sampling to LLM-mediated access challenges established hermeneutic practices. Without critical intervention, “black box” retrieval risks reinforcing algorithmic biases or obscuring the provenance of historical assertions. While current technical best practices focus on computationally optimizing retrieval metrics, this workshop shifts the emphasis to the user’s hermeneutic control. We advocate for a workflow where research goals and the scholar’s interpretation of relevance drive the system, ensuring that computational scalability supports, rather than replaces, deep inquiry. This approach bridges technical RAG implementation with digital hermeneutics, situating AI-driven retrieval and generation within the tradition of critical source criticism.

Workshop Contents

Participants will engage with a custom-built, web-based RAG system (adapted from the SPIEGEL_RAGGED architecture). The system indexes the United Nations General Debate Corpus (UNGDC) 1946–2024, a comprehensive collection of over 10,000 speeches reflecting global political discourse (Jankin et al. 2024).

The workshop follows a structured four-part agenda:

  • Part 0: Setting the Stage (30 min): We survey participants’ current AI usage to establish a baseline, identifying perceived strengths and weaknesses of both machines and scholars. This discussion maps the risks and benefits of augmenting DH workflows with LLMs.
  • Part 1: RAG Foundations (30 min): Instructors provide a conceptual overview of RAG, detailing chunking strategies, embedding models, and vector spaces. The focus is on how technical decisions directly influence utility and interpretation.
  • Part 2: Guided Exploration (30 min): As a plenary group, we investigate a shared research question. Participants verbally direct the workflow (query formulation, retrieval evaluation, and refinement) making collective decisions while navigating the system.
  • Part 3: Hands-On Practice (90 min): Working in small groups, participants pursue questions using the full functionalities of the system including the “LLM-as-a-Judge” method for corpus creation and using LLMs for answering questions on that corpus. Groups must create reflected embedding queries, explicitly define relevance criteria and critique the AI’s selection and generated texts.
  • Part 4: Critical Synthesis (60 min): We finish by revisiting the initial discussion of human-machine capabilities, asking how the workshop experience shifts participants’ understanding of where scholarly authority resides in AI-augmented research.

Didactical Approach

The workshop employs a “learning-by-doing” methodology rooted in the concept of Critical Technical Practice (Agre 1998). Rather than treating the software as a finished product to be mastered, we treat it as an instrument for epistemological reflection. Our pedagogy is structured around three key principles:

  • We meet participants where they are in their own practices. By grounding technical instruction in the participants’ existing insights and experiences, we ensure the RAG tools are understood as augmentations to their specific disciplinary workflows rather than abstract novelties. We emphasize that AI literacy is not just about code, but about understanding the affordances and constraints of the architecture.
  • We break down the “black box” of RAG architecture. By explicitly discussing chunking, embeddings, and vector similarity, we ensure participants understand the mechanical limitations of the system before they use it. This theoretical grounding prevents the anthropomorphising of the AI and grounds subsequent practical work in technical reality.
  • The core of the workshop is the Hands-On testing. Here, didactics focus on using the theoretical knowledge of the embeddings, LLM-as-a-Judge and putting them into practice. This turns the passive search experience into an active construction of a corpus, enforcing explicit articulation of implicit hermeneutic assumptions, then asking questions on that generated corpus to assess the value of LLM text-generation in this context. By working in small groups, participants are forced to verbalize their prompt engineering strategies, revealing that issues are often productive friction points that expose the disconnect between academic nuance and vector representation (Ni et al. 2025).

Diversity & Contribution

This proposal directly addresses the “Remembering” sub-theme by critically engaging with the Memory of the World via the UN corpus. It promotes diversity by making sophisticated data science methodologies accessible to non-programmers, lowering the barrier to entry for complex RAG workflows. The UN corpus, featuring representatives from over 200 countries, offers a diverse, globally representative aggregation of viewpoints on international affairs. Furthermore, we emphasize open science; the workshop code is open-source

https://scm.cms.hu-berlin.de/digital-history/digital-history-forschung/histo_rag/un_ragged. The initial version of the application is running and will be refined through user testing in preparation for the workshop.

and modular, and throughout the workshop we will be introducing participants on how to apply this architecture to their own text collections.

Target Audience & Expected Participants

This workshop is designed for digital humanities scholars, graduate students, and heritage professionals interested in critically engaging with AI technologies in their research practice. No programming experience is required; participants need only a basic familiarity with searching and working with digital text corpora. The workshop is particularly suited for researchers working with digitised historical or archival materials who wish to understand how computational retrieval reshapes traditional hermeneutic practices.

Around 15–20 participants would work well for the workshop. This size allows for effective small-group work while maintaining productive plenary discussions. The hands-on components are designed to accommodate varying levels of technical confidence, with group work fostering peer learning and collective problem-solving.

Technical Requirements

  • Venue: Room with projector/screen and reliable WiFi for up to 20 participants.
  • Participant requirements: Participants must bring their own laptops with a modern web browser.
  • System access: The RAG system is entirely web-based; no software installation or configuration is required. Access credentials will be provided at the workshop, including all API keys for LLM usage.
  • Additional support: Power outlets for participant laptops; whiteboard or flipchart for group discussions would be beneficial but not essential.
References
  1. Agre, P. E. (1998): Toward a Critical Technical Practice: Lessons Learned in Trying to Reform AI, in: Social Science, Technical Systems, and Cooperative Work. Psychology Press.
  2. Gao, Y. / Xiong, Y. / Gao, X. et al. (2024): Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv. http://arxiv.org/abs/2312.10997
  3. Jankin, S. / Baturo, A. / Dasandi, N. (2024): “Words to unite nations: The complete UN General Debate Corpus, 1946–present”, in: Journal of Peace Research. https://doi.org/10.1177/00223433241275335
  4. Lewis, P. / Perez, E. / Piktus, A. / Petroni, F. / Karpukhin, V. / Goyal, N. / Kuttler, H. / Lewis, M. / Yih, W. / Rocktäschel, T. / Riedel, S. / Kiela, D. (2020): “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, in: ArXiv, abs/2005.11401.
  5. Ni, B. / Liu, Z. / Wang, L. et al. (2025): Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey. arXiv. https://doi.org/10.48550/arXiv.2502.06872