Daejeon, July 27–31
The integration of Large Language Models (LLMs) into the digital humanities offers powerful capabilities for a wide spectrum of workflows. This workshop addresses the conference theme of “Engagement” by fostering a critical, active interaction between humanities scholars and AI. We believe that today, when AI agents can generate literature reviews and source corpora with a single click, it is more important than ever that digital humanists understand underlying technologies and display AI literacy. It is then scholar’s responsibility to reflect on how to maintain our own discipline practices when working with AI. Retrieval-Augmented Generation (RAG) explicitly exposes the possibilities and challenges of LLMs, making it the ideal laboratory for developing critical AI literacy. Using a custom RAG system adapted for the United Nations General Debate Corpus and digital humanities practices, this workshop moves beyond “chatting with data.” We aim to teach scholars to work with RAG systems as reflective research aids, ensuring that the “memory of the world” is not just retrieved, but critically interpreted. Participants will leave empowered to maintain their scholarly sovereignty by understanding and critically directing AI systems, rather than being directed by them. This workshop is supported by the NFDI4Memory Methods Innovation Lab.
RAG combines pre-trained parametric memory with non-parametric external knowledge, reducing hallucinations and improving factual consistency (Lewis et al. 2020). While RAG has rapidly evolved to include modular architectures and advanced retrieval strategies (Gao et al. 2024), its application in the humanities often lacks specific epistemological and disciplinary grounding. Recent surveys highlight that while RAG enhances reliability, it introduces new risks regarding explainability and bias that require trustworthy frameworks (Ni et al. 2025).
In historically-focused humanities research, the transition from traditional source sampling to LLM-mediated access challenges established hermeneutic practices. Without critical intervention, “black box” retrieval risks reinforcing algorithmic biases or obscuring the provenance of historical assertions. While current technical best practices focus on computationally optimizing retrieval metrics, this workshop shifts the emphasis to the user’s hermeneutic control. We advocate for a workflow where research goals and the scholar’s interpretation of relevance drive the system, ensuring that computational scalability supports, rather than replaces, deep inquiry. This approach bridges technical RAG implementation with digital hermeneutics, situating AI-driven retrieval and generation within the tradition of critical source criticism.
Participants will engage with a custom-built, web-based RAG system (adapted from the SPIEGEL_RAGGED architecture). The system indexes the United Nations General Debate Corpus (UNGDC) 1946–2024, a comprehensive collection of over 10,000 speeches reflecting global political discourse (Jankin et al. 2024).
The workshop follows a structured four-part agenda:
The workshop employs a “learning-by-doing” methodology rooted in the concept of Critical Technical Practice (Agre 1998). Rather than treating the software as a finished product to be mastered, we treat it as an instrument for epistemological reflection. Our pedagogy is structured around three key principles:
This proposal directly addresses the “Remembering” sub-theme by critically engaging with the Memory of the World via the UN corpus. It promotes diversity by making sophisticated data science methodologies accessible to non-programmers, lowering the barrier to entry for complex RAG workflows. The UN corpus, featuring representatives from over 200 countries, offers a diverse, globally representative aggregation of viewpoints on international affairs. Furthermore, we emphasize open science; the workshop code is open-source https://scm.cms.hu-berlin.de/digital-history/digital-history-forschung/histo_rag/un_ragged. The initial version of the application is running and will be refined through user testing in preparation for the workshop.
This workshop is designed for digital humanities scholars, graduate students, and heritage professionals interested in critically engaging with AI technologies in their research practice. No programming experience is required; participants need only a basic familiarity with searching and working with digital text corpora. The workshop is particularly suited for researchers working with digitised historical or archival materials who wish to understand how computational retrieval reshapes traditional hermeneutic practices.
Around 15–20 participants would work well for the workshop. This size allows for effective small-group work while maintaining productive plenary discussions. The hands-on components are designed to accommodate varying levels of technical confidence, with group work fostering peer learning and collective problem-solving.