Daejeon, July 27–31
The representational capabilities and the advances in ontologies and semantic web technologies that provide enhanced data integration, interoperability, discoverability, contextualisation, and sophisticated querying---particularly in GLAM collections and in effect in the broader Digital Humanities field---are well acknowledged and documented in literature (Hyvönen, 2023; Nappi et al., 2024; Nyhan et al., 2025). However, this potential stand to be unlocked for the Oral History and collective memory domain, with recent efforts focusing on metadata mapping methodologies for the purposes of information exchange (Vrachliotou and Papatheodorou, 2024). Oral history is situated, intersubjective, and performative. It is a record not simply of how what happened is recalled but of how individuals construct meaning from the past (Portelli, 1991). Yet the migration of these living memories into digital environments can create a fundamental epistemological tension when Oral histories encounter schemas and formal representations that strip away nuances and contextualities critical to interpretation, recollection and knowledge production of memory’s sociological significance (Flinn et al., 2024).
Applying event-centric models such as CIDOC-CRM (Doerr, 2003) or foundational ontologies like DOLCE (Gangemi et al., 2002) to oral testimony risks imposing a false objectivity. Such models were designed primarily for cultural heritage objects and entities where catalogue and object-related data tend to be presented as verified and neutral information at the time of computational modelling (but cf. Turner 2020, for example, who problematises such claims from a postcolonial perspective).
Our methodological response to this challenge is to approach and reconceptualise ontology not just as a technical schema for data validation and formal representation but as a hermeneutic instrument that guides computational reading without suppressing the interpretive space. We propose a pathway to Hermeneutic Ontology Engineering that employs LLMs as partners in ontology design, using bottom-up pattern discovery to construct schemas grounded in the actual discourse of the oral testimonies rather than imposed through top-down conceptualisations.
This work is based on oral histories that have been published and publicly available since 2016 under creative commons licenses that permit Creative Commons Attribution–NonCommercial–NoDerivatives (CC BY‑NC‑ND) licences.” This approach is conformant with a CC BY‑NC‑ND licence because it does not produce, disseminate, or substitute the published interviews in adapted, self-contained form. Transcripts are segmented into excerpted narrative units for further analysis by human researchers (who may further incorporate additional digital tools in their analysis of the interviews) and thus function as intermediary research-analytical tools and not replacement interviews. All outputs consist of analytical primitives (ontological relations, classifications, and metadata) that describe interpretive features of the interviews without reproducing or transforming the original interview content, thereby respecting the No-Derivatives restriction. No sensitive or unpublished personal data has been processed.
Beyond this, our hermeneutic framing is itself our methodological response to the datafication concern. Modelling propositions as claims made by specific speakers, respects the oral history principle that testimony retains its speaker and situation, not stand alone as decontextualised, data-driven fact or claim to a singular objective reality.
Digital oral history has predominantly engaged computational methods for access and retrieval rather than interpretation (Smyth et al., 2023). Speech segmentation, automatic transcription, and keyword indexing have enhanced discoverability, but high-resolution semantic modelling is only just emerging. Meanwhile, computational analysis that has been applied to historical and literary corpora has largely relied on aggregative methods aimed at extracting latent patterns from large collections, employing techniques such as topic modelling, sentiment analysis, named entity recognition, and other approaches to information extraction (Jänicke et al., 2015). While powerful for overview, reasoning and enrichment, and potentially beneficial to oral history research, these approaches depend on strong operationalisation choices that compress complex, situated concepts into narrow categorical schemes (Pichler and Reiter, 2022). This can be particularly restrictive for highly subjective oral history testimonies, risking the reduction of heterogeneous voices and contested narratives to apparently stable quantitative patterns that primarily serve semantic alignment and homogenisation.
Recent work has begun testing LLMs specifically on oral history collections targeting the specificities and ambiguities rooted in oral histories. Cherukuri et al. (Cherukuri et al., 2025) demonstrated that prompt-engineered LLMs achieve high accuracy in classifying Japanese American incarceration narratives by sentiment and theme. Widegren (Widegren, 2025) successfully employed LLMs to perform thematic segmentation of oral history transcripts for automatic subject indexing based on the OHMS (Oral History Metadata Synchronizer) schema and controlled vocabulary (Thesaurus for SAMLA and ISOF). Such approaches open new directions to emergent use of Large Language Models (LLMs) as tools for oral history research. Yet, LLMs being probabilistically optimised to produce coherent text, are prone to smooth over irregularities in favour of narrative plausibility, and when applied to fragmentary or contradictory oral testimony, they risk replacing the messy reality of historical formation with fictitious consensus (Ji et al., 2023). In our empirical work on approximately 25 oral history interviews documenting histories of Digital Humanities (Nyhan and Flinn, 2016), we observed that standard LLM extraction workflows systematically hallucinated causal connections and flattened contested claims into apparent facts. We introduce human intervention as a critical step in Hermeneutic Ontology Engineering. By critically reviewing and analysing LLM outputs, we refine the schema to capture the relational structure of remembering, specifically the narrative structures, stances, and sources through which speakers construct memory. Rather than treating extracted propositions as facts, this approach renders the ontology an instrument for making interpretive distinctions visible and queryable.
Hermeneutic Ontology Engineering constitutes the methodological core of the Mixed-methods Digital Oral History (MeDoraH) project, investigating narratives of formation, disruption, and change in the history of Digital Humanities as this is reflected in oral testimonies of domain experts who contributed to the development of the field. A set of 25 interviews was originally created and transcribed between 2012–15, constituting the core material of our investigation and modelling effort, which developed through iterations.
The corpus consists of interviews with well and lesser‑known male and female scholars and technicians who worked in North America and Europe during the formative years of digital humanities. This scope reflects deliberate historical and methodological choices typical of oral history aligned to the linguistic ability and disciplinary expertise of the interviewer (Nyhan), not an implicit claim to represent the field as a whole. Neither the corpus nor the resulting publications have claimed or implied representativeness. On the contrary, titles have been explicitly chosen to signal partiality and contribution rather than completeness, for example Computation and the Humanities: Towards an Oral History of Digital Humanities, where “towards” marks the work as one step within a necessarily larger, plural historiography. We note the role the project has had in raising this awareness of the need for a global oral history of digital humanities and that oral history projects are now being planned or undertaken internationally, for example by Lik Hang Tsui (City University of Hong Kong).
Our early attempts in applying event-centric ontology patterns from cultural heritage modelling encountered fundamental limitations due to oral history’s dialogic interaction and subjective evaluation. Applying schemas and conceptualisations designed for GLAM collections meant either discarding what makes oral history epistemologically distinctive or misrepresenting testimony as verified fact. Instead, we shifted our orientation from reconstructing what happened to modelling how speakers construct accounts of what happened, prioritising interpretive modelling over fact assertion, treating extracted propositions as claims, situated utterances carrying epistemic and evaluative dimensions traceable to specific interview moments.
We adopt a four-layer architecture (Figure 1) maintaining analytic separation between dimensions oral historians treat as epistemologically distinct, forming a cascading order and hierarchy where each layer provides interpretive context for the next.
The Discourse Layer captures the interview as a performative event of interviewer-interviewee conversation. Grounded in Conversation Analysis, it models turn-taking and prompt-response pairs, preserving the intersubjective conditions under which testimony was produced during the interview. All downstream data remains anchored to the interview conversation and speaker attributions.
Figure 1: Four-layer architecture underpinning Hermeneutic Ontology Engineering and the LLM driven workflow.
The Narrative Structure Layer segments discourse into thematically coherent “narrative units”: passages shaped by how speakers configure events into meaningful sequences that capture how recollections are organised into narratives. The Claims Layer atomises narrative units into individual propositions modelled as reified assertions, not as binary facts, providing the required plasticity in the knowledge graph for preserving contested narratives without forcing predetermined resolutions. The Reference Layer contains curated identifiers for actors, organisations, technologies, and events recurring across interviews, providing interoperability for cross-corpus alignment and cross-reference of statements and claims. To better understand the architecture, here is a walkthrough example of an interviewee recalling an early career moment in Figure 2.
To populate this architecture, we developed a Stratified Multi-agent Workflow addressing LLM accuracy degradation on lengthy transcripts (Hong et al., 2025). Rather than processing entire interviews at once, we segment transcripts into narrative units, then extract claims within bounded contexts. The workflow is model-agnostic at the interface level, our results reported here use Google Gemini 2.5 Pro for both extractor and consolidator roles, selected for its instruction following capability on long-context processing and strong performance on structured extraction tasks. Also, open-weight alternatives (Gemma3(gemma-3-27b-it) and Qwen3(Qwen3-30B-A3B-Instruct-2507)) were evaluated for local deployment and workflow reproducibility.
Figure 2: Walk-through of the four-layer architecture applied to a single utterance
We introduce an interactive Workbench (Figure 2) to support human-in-the-loop refinements. The Workbench https://github.com/articoder/MeDoraH_NLP/
Figure 3: Interactive workbench to assist human-in-the-loop refinements
Across 19 interviews, we extracted 176 narrative units containing 1,647 claims. The entity type schema itself emerged through iterative refinement within our Workbench environment. Initial extraction employed schema-free prompting to avoid imposing predetermined categories. However, this produced non-specific and high-level ontological confirmations such as the entity “Concept” which conflated speaker attitudes, academic theories, methodological approaches and disciplinary labels, rendering such entities unusable for meaningful analysis. Using the Workbench’s structural pattern analysis, we examined relation cardinality and entity role distributions across the corpus, revealing that these conflated types occupied systematically different structural positions in the knowledge graph. For example, a person’s attitude towards “Concept” appeared predominantly as a claim-internal stance marker rather than as relational objects (e.g., a person holding a positive view for diversity of approaches in electronic literature). We refined the automated analysis’s 153 initial entity types into an ontological hierarchy of 6 core types (Actor, Event, Artefact, Conceptual Item, Place, Temporal) and 15 subclasses. Our refinement process employed explicit and evaluative stance attributes within the Claims Layer, separating attitudinal content from referenced entities, and structure-aware clustering for distinguishing entity types. During the refinement process we validated against textual evidence using the Workbench’s provenance tracking. Both refined ontology and the Workbench are available in GitHub.
Evaluating the hermeneutic framework presents distinct challenges: our goal is not to maximise extraction recall against a fixed gold standard but to assess whether the ontological schema supports meaningful interpretive distinctions. We employ three complementary evaluation strategies. First, provenance coverage: we measure the proportion of extracted claims that maintain traceable links to source transcript spans. Second, multi-pass examination: we assess convergence and divergence across the three extraction configurations, treating disagreement not as error but as a diagnostic signal for interpretive ambiguity. Third, expert validation: domain specialists review sampled claims and ontological categories against their close reading of the source material. A hand-coded dataset of 20 interviews produced by a historian on the project team serves as a qualitative baseline for comparing entity coverage and relational patterns.
Our work contributes a hermeneutic, evidence-first approach to computational modelling of oral history where ontological commitments function as interpretive hypotheses rather than truth assertions. Bottom-up schema induction, coupled with bounded-context extraction and provenance anchoring, substantially improved quality while preserving contestation and traceability. The resulting ontology increased analytic resolution for downstream inquiry, and the framework supports memory-orientated interpretation by making distinctions between speaking, narrating, claiming, and referencing explicit and queryable.