Daejeon, July 27–31
The data turn in the humanities (Wenqi Li et al. 2024; Zlodi 2023) has transformed knowledge production and dissemination. Yet it has exposed structural weaknesses: fragmentation, lack of standardisation, limited accessibility, and uneven long-term preservation (Lee / Chung 2023; Min et al. 2023). At the University of Barcelona (UB), a diagnostic study revealed the same problem: dozens of humanities databases from projects, theses, and research initiatives are dispersed across servers, formats, and documentation standards, limiting their scholarly and societal value. This mirrors international concerns about the vulnerability of humanities data without robust institutional infrastructures (Knight et al. 2024).
The Digital Humanities Hub at the UB (the Hub) is an institutional response to this condition of abundance without infrastructure. It reframes humanities datasets as shared institutional assets requiring coordinated governance, stewardship, and reuse, and responds to concerns that sustainability must be understood not only technically but also in relation to public value, social engagement, and cultural stewardship (Gómez Zapata 2021; Kairiss et al. 2023).
The paper advances a model organised around three pillars: standardisation aligned with FAIR principles —Findable, Accessible, Interoperable, and Reusable— and Linked Open Data (LOD); a knowledge-graph-based semantic layer combining neurosymbolic artificial intelligence (AI) and retrieval-augmented generation for explainable, bias-aware access; and participatory, human-centred design. Within this framework, it makes four contributions: it reframes fragmented humanities databases as governed institutional assets; outlines a FAIR/LOD-to-Knowledge Graph (KG) integration pathway; proposes a transparent retrieval model grounded in neurosymbolic AI; and adapts Social Return on Investment (SROI) to assess the institutional and societal value of digital humanities infrastructures.
The UB diagnostic confirms Samra et al.’s (2021) “first-mile problem” in data management: while many projects generate digital data, few develop workflows for long-term curation, integration, or reuse. UB databases suffer from duplicated entities, inconsistent metadata, heterogeneous formats, and poor documentation—issues also noted by Graefe et al. (2025). Without structured preservation, many datasets risk abandonment once funding ends, reflecting patterns highlighted by Knight et al. (2024).
Moreover, as Herrera-Cubides et al. (2023) argue, sustainability challenges are exacerbated by unclear licensing, inconsistent storage environments, and the absence of institutional policies. The UB study confirms this vulnerability: data is fragile, underused, and difficult to access. This is aggravated by the underrepresentation of women and minoritised groups rooted in historical production biases, an issue emphasised in digital humanities and AI bias research (Finzel 2025; Gomaa / Feld 2023). The Hub addresses this by reframing humanities data as institutional assets requiring coordinated governance, semantic alignment, and stewardship.
The first pillar implements three phases: extraction and localisation, reflecting emphasis on clear inventories to avoid data loss (Zlodi 2023; Min et al. 2023); refinement and normalisation, following best practices for metadata coherence and cleaning alongside standards such as Dublin Core and Text Encoding Initiative (TEI) (Graefe et al. 2025; Brahmia et al. 2022; Zlodi 2023); and integration of FAIR and LOD principles for interoperability and reuse (Candela 2023; Wilkinson et al. 2016; Zlodi 2023).
The transition to RDF and SPARQL infrastructures (Candela 2023; Gao et al. 2024) places UB within a global network of interoperable heritage data. CIDOC-CRM, identified by Antonini et al. (2023) as essential for cultural heritage, provides the basis for modelling relationships. This responds to critiques that without standardisation and FAIR alignment, data remains invisible, irreusable, and institutionally inefficient (Graefe et al. 2025; Lee / Chung 2023).
Recent work in neurosymbolic artificial intelligence (Finzel, 2025; Mileo, 2025) shows that combining symbolic reasoning (ontologies, rules, entity relationships) with neural architectures (Large Language Models, embeddings) provides a framework for transparent, bias-aware AI systems. The Hub builds on these findings through a multi-layered knowledge graph integrating an ontology layer informed by CIDOC-CRM, Dublin Core, and domain schemas (Antonini et al. 2023; Zlodi 2023); an entities layer harmonising data from UB projects, institutional collections, and open-data sources such as Wikidata; and a reasoning layer enabling inferencing and semantic constraint.
A tabular dataset extracted from judicial case files concerning individuals prosecuted under the Ley de Vagos y Maleantes (“Law on Vagrants and Malefactors”), as amended by the Franco regime on 15 July 1954 to extend its repressive scope to homosexuals, or a graph database derived from Francoist war records, may initially exist as an isolated doctoral research resource. Once normalised and mapped to a shared semantic layer, these materials can be connected to persons, places, proceedings, institutional actors, and documentary sources. A query about the relation between a defendant, a legal process, and recurring actors can then be answered through retrieval grounded in the knowledge graph, with the response traceable to explicit records and semantic relations.
This architecture reflects Gao et al. (2024) and Santini’s (2024) emphasis on knowledge graphs for large-scale, semantically rich information environments in dialogue with machine learning. A Retrieval-Augmented Generation (RAG) system grounds responses in the curated KG, mitigating hallucinations and systemic biases (Finzel 2025; Gomaa / Feld 2023). This moves UB “from silos to semantic ecosystems”, enabling discovery, cross-database exploration, and analysis.
Drawing on human-centred AI frameworks (Gomaa / Feld 2023; Schmager et al. 2023), the Hub embeds participatory design through co-design workshops, scenario-based design and card-sorting, community validation of KG entities (e.g. Gardasevic / Lamba 2024), and integration with Wikimedia ecosystems. These activities follow an iterative cycle in which stakeholder workshops inform prototypes, later refined through evaluation and validation before integration into the Hub’s workflows. This reflects Kairiss et al.’s (2023) claim that cultural infrastructures must integrate social and community dimensions for sustainability and Gómez Zapata’s (2021) emphasis on public value and civic engagement.
The UB diagnostic proposes routes for institutional improvement. The Hub incorporates these proposals through an inventory of digital assets, a standardisation framework, and a SROI perspective. The first result is a map of humanities databases at UB, reflecting the principle that visibility is the condition for preservation and reuse (Samra et al. 2021; Zlodi 2023).
This model is particularly important for small-scale humanities databases created in doctoral research, pilot projects, or short-term initiatives. Such resources often contain unique manually extracted data—for example, from judicial archives or specialised documentary corpora—yet remain difficult to discover, compare, or reuse beyond their original setting. By integrating them into a shared semantic environment, the Hub increases their visibility, supports connections with related datasets, and creates conditions for broader scholarly and public reuse.
The three-phase workflow—inventory, refinement, FAIR/LOD integration—operates within the Hub, responding to abandoned projects, inconsistent formats, lack of persistent identifiers, and insufficient metadata (Graefe et al. 2025; Lee / Chung 2023). In practical terms, resources are identified and documented, cleaned and aligned through shared metadata and documentation practices, and mapped to a semantic layer that supports knowledge-graph integration. On this basis, the Hub enables retrieval, cross-dataset discovery, and public-facing access, while creating conditions for evaluation through reuse, preservation, and social-impact indicators.
A key contribution of the UB study is a qualitative SROI framework that extends beyond financial metrics (Mengsiying Li / Wang 2024). In the Hub, this framework operates through four dimensions: economic indicators such as reduced duplication and more efficient reuse; conservation indicators such as stabilising at-risk datasets and improving documentation; scientific indicators such as dataset reuse across projects and interdisciplinary connections; and social indicators such as educational uptake, public engagement, visibility of underrepresented communities, and collective memory through accessible digital infrastructures, in line with understandings of social value in cultural institutions (Gómez Zapata 2021). Together, these indicators allow the Hub to assess whether data has been preserved and has generated institutional and public value.
The Hub embodies values central to engaged digital humanities: it creates open, inclusive infrastructures; fosters collaboration across academia, memory institutions, and society; supports public-facing pedagogies that enable students and citizens to contribute to knowledge graphs; supports cultural memory and the representation of historically marginalised groups; and treats digital research infrastructure as critical public infrastructure (Kairiss et al. 2023). It prioritises datasets documenting underrepresented communities, supports multilingual access, and incorporates bias-aware auditing of corpora used for AI services. This is particularly relevant for datasets documenting historically marginalised subjects, including women, politically repressed individuals, and populations criminalised under discriminatory legal frameworks, whose traces often remain dispersed across fragile or project-bound research resources.
The Digital Humanities Hub at the UB proposes a sustainable, socially responsible model for humanities data in institutional contexts. By combining FAIR and LOD standards (Candela 2023; Zlodi 2023), knowledge-graph architectures (Gao et al. 2024), neurosymbolic AI (Finzel, 2025; Mileo, 2025), and participatory methodologies, it turns dispersed datasets into a cohesive and open knowledge ecosystem. Grounded in institutional analysis and international literature, the Hub provides the UB with a framework for consolidating and extending its digital humanities infrastructure. It links technical integration, public engagement, and stewardship, while positioning the institution to play a leading role in socially responsible digital humanities infrastructures.