Daejeon, July 27–31
This short paper presents ongoing research from a thesis project on cultural data interoperability. It focuses on the effect of classification and categorization on historical implicit knowledge, while exploring best practices for building judicious and equitable digital humanities (DH) projects and datasets. In line with the DH2026 themes of “Remembering” and “Annotating”, the study examines how to model tacit knowledge about women and their labor inside archaeological records. An equity-aware approach can reshape how human historical knowledge is expressed and shared across cultural boundaries, whether these boundaries arise nationally, linguistically, or technically (ex: semantic web, institutional databases).
Existing research on semantic interoperability in cultural heritage has demonstrated the potential of ontologies and linked open data (LOD) to connect collections, and highlighted the challenges of aligning complex descriptions with semantic models (Du Chateau et al. 2020; Liu 2020; Ore / Eide 2009). Additionally, the application of FAIR (Findable, Accessible, Interoperable, Reusable) principles to the Humanities has shown that data management practices risk losing contextual detail and provenance. This could happen especially when complex archival knowledge is compressed into standardised metadata for reuse (Kansa et al. 2025). For example, field notes describing occupation probabilities, and culturally local knowledge and practices can be flattened, even removed while modeling knowledge. They’re mapped with generic descriptions (like types in CIDOC CRM) that obscure uncertainty, temporality, or in E53_Place case, how well fine‑grained senses of locality are captured (Eide 2025; Metilli et al. 2026). This paper addresses that gap by using a set of archaeological records, as a testbed, for assessing and testing how well CIDOC CRM (an event-centric structure) can represent tacit knowledge across historically layered collections. This work builds on heritage interoperability efforts, such as the Archaeology Data Service uses ARIADNE and associated controlled vocabularies, or Beyond Notability uses a bespoke LOD model.
This study draws on several archival collections held in Canadian university archives, documenting archaeological fieldwork and heritage management activities in Ontario, Nunavut, Yukon, England, and western Canada. These documents include reports, excavation permits, shovel test records, daily notes, surveyors’ documentation, and extensive photographic materials capturing landscapes, artefacts, and field teams (the final scope is yet to be determined and hinges on access to collections). Taken together, they record sites and objects, but also research practices, collaborations, and relationships with the land and communities over several decades. These records provide enough information so we may ask how a CIDOC CRM-oriented model can represent people, places, events, and objects in these documents in ways that respect their historical and cultural specificity while enabling cross-structure interoperability? The project investigates the limits encountered when encoding and assigning persistent identifiers to entities drawn from historically layered archival descriptions; and the extent to which categorization can accommodate specific concepts to specific times, languages, and cultures within a broadly shared ontology.
The methodological approach combines qualitative data modeling with practical encoding and conversion workflows. First, selected people, places, events, and material objects are encoded in TEI, starting from existing finding aids and representative documents. This step foregrounds the interpretive work involved in turning narrative description into structured markup, including decisions about what to segment as an entity (based on salience, recurrence across documents, or need to portray) and how to represent uncertainty or contestation. Second, each entity is linked to, or assigned, a persistent identifier, drawing on existing authority files and external LOD sources where possible, and creating local URIs when necessary. This exposes tensions between local, archival naming practices and global authority infrastructures, especially for places and persons underrepresented in existing registries (Tóth-Czifra 2020). For example, places, persons or culturally defined practices (activities) may require inventing custom role classes; they’d not be interoperable within existing systems but would challenge the methods that make us discriminate against certain kinds of knowledge.
The analysis focuses on two intertwined dimensions: semantic adequacy and interpretive/ethical implications. On the semantic side, the paper discusses examples where CIDOC CRM provides a strong fit (E7_Activity for modeling excavation events, artefact production, documentation processes, and participatory events) and cases where it struggles (ex: collaborative roles that exceed standard E39_Actor categories or nuanced distinctions in field practices), revealing limits in representing situated labor practices. On the interpretive and ethical side, the study adopts a feminist and critical perspective to examine how modeling choices inherited from archival descriptions, and authority structures shape who and what becomes visible (D’Ignazio / Klein 2020; Noble 2018).
Findings suggest that CIDOC CRM and its extensions can significantly improve interoperability across the documents, making it possible to query all events involving a given person across regions, or to connect photographs, reports, and permits related to a specific site. But some historically significant distinctions, important to archaeologists and communities, are difficult to encode without resorting to local extensions, and authority-based identifiers often reproduce existing asymmetries in whose names and places are standardized. And if on their side, persistent identifiers enable cross-collection connections, they remain hardly applicable where authority records are absent or biased. The study also shows that steps toward FAIR data (such as creating persistent identifiers and publishing in RDF) do not automatically guarantee reusability if legal, ethical, and community considerations remain unresolved, or if the translation from archival description to ontology erases context (among others).
These observations resonate with wider concerns about the risk of losing thick description and the dependency of researchers on external infrastructures and policies. This short paper offers both a practical workflow and a critical reflection. Practically, it provides a reusable sequence (from TEI encoding to CIDOC CRM-based RDF, and graph visualisation further down the road) for enriching and connecting small, heterogeneous archival collections broadly. Critically, it argues that semantic interoperability for cultural heritage cannot be evaluated only in technical terms but must also be assessed in relation to whose memories, places are made legible in machine-readable format. Because standardized decisions privilege standardized mappings that can affect whether certain forms of knowledge or labor will be visible in interoperable systems (Bowker / Star 1999; Costanza-Chock 2020).