Daejeon, July 27–31
The Icelandic Centre for Digital Humanities and Arts (DARIAH-IS/CDHA) operates as a national hub for digital research infrastructure in the humanities and arts, coordinating tools, software, databases, and expertise. As a collaboration between fifteen institutions, the Centre’s mission is to lower technical and institutional barriers to digital research, while supporting meaningful public engagement with cultural heritage. In this national ecosystem, the CDHA Web Portal is conceived as a core piece of connective infrastructure rather than as a single, monolithic database. Despite Iceland’s relatively small population, its digital cultural heritage landscape is remarkably rich yet structurally fragmented. Legacy databases run on bespoke systems with limited or no APIs, newer platforms follow divergent metadata practices and vocabulary standards, and access regimes range from fully open public collections to tightly restricted scholarly and broadcasting archives. Researchers, students, and the general public consequently face high transaction costs when attempting to discover materials across institutions or media types.
This paper presents the technical and pedagogical design of the new CDHA Web Portal, which aims to transform this fragmented landscape into a unified point of discovery. The portal is explicitly framed around “engagement” rather than simple aggregation: it is designed not only to make records findable, but to activate them for interdisciplinary scholarship, teaching, and public participation.
The challenge of aggregating heterogeneous cultural heritage data is well established in the Digital Humanities and GLAM sectors. Large-scale initiatives such as Europeana have demonstrated the power of Linked Open Data, shared schemas, and semantic interoperability at continental scale (Kail 2011), but have also revealed the costs of strict top-down alignment to complex ontologies such as CIDOC-CRM (cf. Doerr et al. 2007). For smaller institutions maintaining legacy systems or operating with limited technical capacity, full data migration into a shared model can be financially and organizationally prohibitive. Rather than treating ontological harmonization as a precondition for participation, the portal adopts a pragmatic, index-first strategy that emphasises inclusion, incremental improvement, and sustainability for smaller actors.
The core purpose of the portal is to enable users to identify and traverse connections between the consortium’s databases without requiring those databases to be structurally homogenized in advance. To achieve this, the system constructs a unified search index using a free-standing, typo-tolerant engine optimised for speed and user experience that stores a “minimum set” of harmonised metadata fields: title, creator, date, place, related people, and a small number of type or subject descriptors. These fields are harvested from institutional databases via REST APIs where available, or through scheduled ETL workflows for legacy SQL/FileMaker systems and other non-API environments. This approach allows “static” or intermittently updated collections to be represented alongside more dynamic, API-driven repositories, while keeping the authoritative records and detailed metadata in their original systems.
Methodologically, the portal is designed as an overlay rather than a replacement for institutional infrastructures. Each indexed record points users back to the underlying system and, when necessary, to the physical location of non-digitised material, foregrounding custodianship and local archival practices. This model reduces the pressure on partners to undertake costly, one-off data migrations and instead encourages gradual enhancement of the exposed “minimum set” over time. At the same time, it allows the portal team to focus on interface design, relevance ranking, and cross-collection services rather than on maintaining yet another full-scale collection management system.
A key technical and epistemological challenge lies in reconciling divergent controlled vocabularies and descriptive practices. Instead of enforcing a single top-down ontology, the portal implements a “semantic grouping” layer that links domain-specific vocabularies into a flexible superset. General users can search using broad concepts and receive results mapped from different institutional terminologies, while expert users retain access to more granular, domain-specific terms through faceted search and advanced filters.
Beyond search, the Web Portal is conceived as a pedagogical and research workbench. The interface includes tools for building custom collections that cut across institutions, time periods, and media types, drawing inspiration from existing Icelandic systems for museums and libraries, such as the Sarpur system, while extending them into a cross-domain environment. Students, teachers, and researchers can assemble thematic corpora, combining museum objects, archival records, audio-visual broadcasts, and manuscript descriptions. These curated sets can be exported to external environments for computational analysis or explored within the portal using integrated temporal and spatial visualizations. By embedding these tools into the core design, the portal supports a “loop” of engagement in which users do not merely locate items but actively reorganize and reinterpret them. Classroom activities can involve building and comparing alternative collections, tracing how different institutions narrate the same events, or juxtaposing digitised and non-digitised material through metadata. This design aligns with the conference theme by treating infrastructure not as a neutral pipeline but as a site where methodological reflection, interpretive practice, and technological critique can be foregrounded.
The project also seeks to support diversity and accessibility in several interrelated senses. First, it bridges the gap between well-resourced university projects and smaller, underfunded heritage institutions by offering a low-barrier path into a shared discovery environment. Second, it accommodates a wide range of media in a manner that makes them jointly visible without flattening their differences. Third, it addresses complex copyright and rights-management regimes by treating metadata as a first-class citizen: restricted materials can still appear in search results with descriptive records, preventing legal constraints from rendering significant parts of cultural history effectively invisible to researchers and the public. For a small-language community like Icelandic, this approach also has implications for multilingual engagement and AI-mediated access. By providing a rich, structured, and semantically linked index over heterogeneous collections, the portal creates a foundation on which future translation and recommendation services can operate in a responsible way, while keeping human curatorial practices and local knowledge at the centre. Rather than retrofitting engagement onto an existing technical stack, the project treats engagement as a primary design principle.
The paper will report on the first implementation cycle of this model, including architectural decisions, indexing and vocabulary strategies, and early use cases from teaching and research. It will argue that such an approach offers a transferable blueprint for other small nations and linguistic communities seeking to build sustainable, inclusive, and critically informed digital heritage infrastructures.