DH 2026

Daejeon, July 27–31

Fri, July 3109:00–10:30S021105
Long Paper

Digital Humanities Hub: From Research Silos to Knowledge Graphs and Neurosymbolic AI

Núria Ferran-Ferran
Universitat de Barcelona, Spain · nferranf@ub.edu
Miquel Centelles Velilla
Universitat de Barcelona, Spain · miquel.centelles@ub.edu

Introduction

The data turn in the humanities (Wenqi Li et al. 2024; Zlodi 2023) has transformed knowledge production and dissemination. Yet it has exposed structural weaknesses: fragmentation, lack of standardisation, limited accessibility, and uneven long-term preservation (Lee / Chung 2023; Min et al. 2023). At the University of Barcelona (UB), a diagnostic study revealed the same problem: dozens of humanities databases from projects, theses, and research initiatives are dispersed across servers, formats, and documentation standards, limiting their scholarly and societal value. This mirrors international concerns about the vulnerability of humanities data without robust institutional infrastructures (Knight et al. 2024).

The Digital Humanities Hub at the UB (the Hub) is an institutional response to this condition of abundance without infrastructure. It reframes humanities datasets as shared institutional assets requiring coordinated governance, stewardship, and reuse, and responds to concerns that sustainability must be understood not only technically but also in relation to public value, social engagement, and cultural stewardship (Gómez Zapata 2021; Kairiss et al. 2023).

The paper advances a model organised around three pillars: standardisation aligned with FAIR principles —Findable, Accessible, Interoperable, and Reusable— and Linked Open Data (LOD); a knowledge-graph-based semantic layer combining neurosymbolic artificial intelligence (AI) and retrieval-augmented generation for explainable, bias-aware access; and participatory, human-centred design. Within this framework, it makes four contributions: it reframes fragmented humanities databases as governed institutional assets; outlines a FAIR/LOD-to-Knowledge Graph (KG) integration pathway; proposes a transparent retrieval model grounded in neurosymbolic AI; and adapts Social Return on Investment (SROI) to assess the institutional and societal value of digital humanities infrastructures.

The Problem: Abundance Without Infrastructure

The UB diagnostic confirms Samra et al.’s (2021) “first-mile problem” in data management: while many projects generate digital data, few develop workflows for long-term curation, integration, or reuse. UB databases suffer from duplicated entities, inconsistent metadata, heterogeneous formats, and poor documentation—issues also noted by Graefe et al. (2025). Without structured preservation, many datasets risk abandonment once funding ends, reflecting patterns highlighted by Knight et al. (2024).

Moreover, as Herrera-Cubides et al. (2023) argue, sustainability challenges are exacerbated by unclear licensing, inconsistent storage environments, and the absence of institutional policies. The UB study confirms this vulnerability: data is fragile, underused, and difficult to access. This is aggravated by the underrepresentation of women and minoritised groups rooted in historical production biases, an issue emphasised in digital humanities and AI bias research (Finzel 2025; Gomaa / Feld 2023). The Hub addresses this by reframing humanities data as institutional assets requiring coordinated governance, semantic alignment, and stewardship.

Pillar 1: Standardisation and FAIR/LOD Integration

The first pillar implements three phases: extraction and localisation, reflecting emphasis on clear inventories to avoid data loss (Zlodi 2023; Min et al. 2023); refinement and normalisation, following best practices for metadata coherence and cleaning alongside standards such as Dublin Core and Text Encoding Initiative (TEI) (Graefe et al. 2025; Brahmia et al. 2022; Zlodi 2023); and integration of FAIR and LOD principles for interoperability and reuse (Candela 2023; Wilkinson et al. 2016; Zlodi 2023).

The transition to RDF and SPARQL infrastructures (Candela 2023; Gao et al. 2024) places UB within a global network of interoperable heritage data. CIDOC-CRM, identified by Antonini et al. (2023) as essential for cultural heritage, provides the basis for modelling relationships. This responds to critiques that without standardisation and FAIR alignment, data remains invisible, irreusable, and institutionally inefficient (Graefe et al. 2025; Lee / Chung 2023).

Pillar 2: Knowledge Graphs and Neurosymbolic AI

Recent work in neurosymbolic artificial intelligence (Finzel, 2025; Mileo, 2025) shows that combining symbolic reasoning (ontologies, rules, entity relationships) with neural architectures (Large Language Models, embeddings) provides a framework for transparent, bias-aware AI systems. The Hub builds on these findings through a multi-layered knowledge graph integrating an ontology layer informed by CIDOC-CRM, Dublin Core, and domain schemas (Antonini et al. 2023; Zlodi 2023); an entities layer harmonising data from UB projects, institutional collections, and open-data sources such as Wikidata; and a reasoning layer enabling inferencing and semantic constraint.

A tabular dataset extracted from judicial case files concerning individuals prosecuted under the Ley de Vagos y Maleantes (“Law on Vagrants and Malefactors”), as amended by the Franco regime on 15 July 1954 to extend its repressive scope to homosexuals, or a graph database derived from Francoist war records, may initially exist as an isolated doctoral research resource. Once normalised and mapped to a shared semantic layer, these materials can be connected to persons, places, proceedings, institutional actors, and documentary sources. A query about the relation between a defendant, a legal process, and recurring actors can then be answered through retrieval grounded in the knowledge graph, with the response traceable to explicit records and semantic relations.

This architecture reflects Gao et al. (2024) and Santini’s (2024) emphasis on knowledge graphs for large-scale, semantically rich information environments in dialogue with machine learning. A Retrieval-Augmented Generation (RAG) system grounds responses in the curated KG, mitigating hallucinations and systemic biases (Finzel 2025; Gomaa / Feld 2023). This moves UB “from silos to semantic ecosystems”, enabling discovery, cross-database exploration, and analysis.

Pillar 3: Participatory and Human-Centred Design

Drawing on human-centred AI frameworks (Gomaa / Feld 2023; Schmager et al. 2023), the Hub embeds participatory design through co-design workshops, scenario-based design and card-sorting, community validation of KG entities (e.g. Gardasevic / Lamba 2024), and integration with Wikimedia ecosystems. These activities follow an iterative cycle in which stakeholder workshops inform prototypes, later refined through evaluation and validation before integration into the Hub’s workflows. This reflects Kairiss et al.’s (2023) claim that cultural infrastructures must integrate social and community dimensions for sustainability and Gómez Zapata’s (2021) emphasis on public value and civic engagement.

Institutional Impact: Integrating Results from the UB Study

The UB diagnostic proposes routes for institutional improvement. The Hub incorporates these proposals through an inventory of digital assets, a standardisation framework, and a SROI perspective. The first result is a map of humanities databases at UB, reflecting the principle that visibility is the condition for preservation and reuse (Samra et al. 2021; Zlodi 2023).

This model is particularly important for small-scale humanities databases created in doctoral research, pilot projects, or short-term initiatives. Such resources often contain unique manually extracted data—for example, from judicial archives or specialised documentary corpora—yet remain difficult to discover, compare, or reuse beyond their original setting. By integrating them into a shared semantic environment, the Hub increases their visibility, supports connections with related datasets, and creates conditions for broader scholarly and public reuse.

The three-phase workflow—inventory, refinement, FAIR/LOD integration—operates within the Hub, responding to abandoned projects, inconsistent formats, lack of persistent identifiers, and insufficient metadata (Graefe et al. 2025; Lee / Chung 2023). In practical terms, resources are identified and documented, cleaned and aligned through shared metadata and documentation practices, and mapped to a semantic layer that supports knowledge-graph integration. On this basis, the Hub enables retrieval, cross-dataset discovery, and public-facing access, while creating conditions for evaluation through reuse, preservation, and social-impact indicators.

A key contribution of the UB study is a qualitative SROI framework that extends beyond financial metrics (Mengsiying Li / Wang 2024). In the Hub, this framework operates through four dimensions: economic indicators such as reduced duplication and more efficient reuse; conservation indicators such as stabilising at-risk datasets and improving documentation; scientific indicators such as dataset reuse across projects and interdisciplinary connections; and social indicators such as educational uptake, public engagement, visibility of underrepresented communities, and collective memory through accessible digital infrastructures, in line with understandings of social value in cultural institutions (Gómez Zapata 2021). Together, these indicators allow the Hub to assess whether data has been preserved and has generated institutional and public value.

Contribution to Public-Facing Digital Humanities

The Hub embodies values central to engaged digital humanities: it creates open, inclusive infrastructures; fosters collaboration across academia, memory institutions, and society; supports public-facing pedagogies that enable students and citizens to contribute to knowledge graphs; supports cultural memory and the representation of historically marginalised groups; and treats digital research infrastructure as critical public infrastructure (Kairiss et al. 2023). It prioritises datasets documenting underrepresented communities, supports multilingual access, and incorporates bias-aware auditing of corpora used for AI services. This is particularly relevant for datasets documenting historically marginalised subjects, including women, politically repressed individuals, and populations criminalised under discriminatory legal frameworks, whose traces often remain dispersed across fragile or project-bound research resources.

Conclusion

The Digital Humanities Hub at the UB proposes a sustainable, socially responsible model for humanities data in institutional contexts. By combining FAIR and LOD standards (Candela 2023; Zlodi 2023), knowledge-graph architectures (Gao et al. 2024), neurosymbolic AI (Finzel, 2025; Mileo, 2025), and participatory methodologies, it turns dispersed datasets into a cohesive and open knowledge ecosystem. Grounded in institutional analysis and international literature, the Hub provides the UB with a framework for consolidating and extending its digital humanities infrastructure. It links technical integration, public engagement, and stewardship, while positioning the institution to play a leading role in socially responsible digital humanities infrastructures.

References
  1. Antonini, Alessio / Adamou, Alessandro / Suárez-Figueroa, Mari Carmen / Benatti, Francesca (2023): “Experiential Observations: An Ontology Pattern-Based Study on Capturing the Potential Content within Evidences of Experiences”, in: Journal on Computing and Cultural Heritage 16, 3: 1–30. DOI: 10.1145/3586078.
  2. Brahmia, Zouhaier / Hamrouni, Hind / Bouaziz, Rafik (2022): “TempoX: A disciplined approach for data management in multi-temporal and multi-schema-version XML databases”, in: Journal of King Saud University - Computer and Information Sciences 34, 1: 1472–1488. DOI: 10.1016/j.jksuci.2019.08.009.
  3. Candela, Gustavo (2023): “An automatic data quality approach to assess semantic data from cultural heritage institutions”, in: Journal of the Association for Information Science and Technology 74, 7: 866–878. DOI: 10.1002/asi.24761.
  4. Finzel, Bettina (2025): “Toward trustworthy AI with integrative explainable AI frameworks”, in: It - Information Technology 67, 1: 20–45. DOI: 10.1515/itit-2025-0007.
  5. Gao, Yan / Xiong, Guanyu / Li, Haijiang / Richards, Jarrod (2024): “Exploring bridge maintenance knowledge graph by leveraging GrapshSAGE and text encoding”, in: Automation in Construction 166: 105634. DOI: 10.1016/j.autcon.2024.105634.
  6. Gardasevic, Stanislava / Lamba, Manika (2024): “Knowledge Graph for Academic Libraries: A User-Centered Design Approach” <https://www.ideals.illinois.edu/items/130413> [02.05.2026].
  7. Gomaa, Amr / Feld, Michael (2023): “Towards Adaptive User-centered Neuro-symbolic Learning for Multimodal Interaction with Autonomous Systems”, in: International Conference on Multimodal Interaction: 689–694. DOI: 10.1145/3577190.3616121.
  8. Gómez Zapata, Jonathan Daniel (2021): Las bibliotecas tienen valor: Una mirada desde el retorno social de la inversión. Medellín: Alcaldía.
  9. Graefe, Adam S. L. / Hübner, Miriam R. / Rehburg, Filip / Sander, Steffen / Klopfenstein, Sophie A. I. / Alkarkoukly, Samer / Grönke, Ana / Weyersberg, Annic / Danis, Daniel / Zschüntzsch, Jana / Nyoungui, Elisabeth F. / Wiegand, Susanna / Kühnen, Peter / Robinson, Peter N. / Beyan, Oya / Thun, Sylvia (2025): “An ontology-based rare disease common data model harmonising international registries, FHIR, and Phenopackets”, in: Scientific Data 12, 1: 234. DOI: 10.1038/s41597-025-04558-z.
  10. Herrera-Cubides, Jhon Francined / Gaona-García, Paulo Alonso / Montenegro-Marin, Carlos Enrique / Sánchez-Alonso, Salvador (2023): “The Relevance of Open Data Principles for the Web of Data”, in: Journal of Electrical and Computer Engineering 2023: 1–17. DOI: 10.1155/2023/4854965.
  11. Kairiss, Andris / Geipele, Ineta / Olevska-Kairisa, Irina (2023): “Sustainability of Cultural Heritage-Related Projects: Use of Socio-Economic Indicators in Latvia”, in: Sustainability 15, 13: 10109. DOI: 10.3390/su151310109.
  12. Knight, Dawn / Khallaf, Nouran / Rayson, Paul / El-Haj, Mahmoud / Ezeani, Ignatius / Morris, Steve (2024): “FreeTxt: A corpus-based bilingual free-text survey and questionnaire data analysis toolkit”, in: Applied Corpus Linguistics 4, 3: 100103. DOI: 10.1016/j.acorp.2024.100103.
  13. Lee, Jungyeoun / Chung, Eunkyung (2023): “Research information service development plan based on an analysis of the digital scholarship lifecycle experience of humanities scholars in Korea: A qualitative study”, in: Science Editing 10, 2: 127–134. DOI: 10.6087/kcse.309.
  14. Li, Mengsiying / Wang, Tai (2024): “Optimizing learning return on investment: Identifying learning strategies based on user behavior characteristic in language learning applications”, in: Education and Information Technologies 29, 6: 6651–6681. DOI: 10.1007/s10639-023-12078-9.
  15. Li, Wenqi / Zhang, Pengyi / Wang, Jun (2024): “Analysing humanities scholars’ data seeking behaviour patterns using Ellis’ model”, in: Information Research an International Electronic Journal 29, 2: 401–418. DOI: 10.47989/ir292835.
  16. Mileo, Alessandra (2025): “Towards a neuro-symbolic cycle for human-centered explainability”, in: Neurosymbolic Artificial Intelligence 1: NAI-240740. DOI: 10.3233/NAI-240740.
  17. Min, Kyoungju / Cho, Jeongyun / Jung, Manho / Lee, Hyangbae (2023): “Analysis of Impact Between Data Analysis Performance and Database”, in: Journal of Information and Communication Convergence Engineering 21, 3: 244–251. DOI: 10.56977/jicce.2023.21.3.244.
  18. Samra, Halima / Li, Alice / Soh, Ben (2021): “Design of a clinical database to support research purposes: Challenges and solutions”, in: International Journal of Advanced and Applied Sciences 8, 3: 21–29. DOI: 10.21833/ijaas.2021.03.003.
  19. Santini, Cristian (2024): “Combining language models for knowledge extraction from Italian TEI editions”, in: Frontiers in Computer Science 6: 1472512. DOI: 10.3389/fcomp.2024.1472512.
  20. Schmager, Stefan / Pappas, Ilias / Vassilakopoulou, Polyxeni (2023): Defining human-centered AI: A comprehensive review of HCAI literature <https://aisel.aisnet.org/mcis2023/13/> [02.05.2026].
  21. Wilkinson, Mark D. / Dumontier, Michel / Aalbersberg, IJsbrand Jan / Appleton, Gabrielle / Axton, Myles / Baak, Arie / Blomberg, Niklas / Boiten, Jan-Willem / da Silva Santos, Luiz Bonino / Bourne, Philip E. (2016): “The FAIR Guiding Principles for scientific data management and stewardship”, in: Scientific Data 3, 1: 1–9. DOI: 10.1038/sdata.2016.18.
  22. Zlodi, Goran (2023): “Information and Documentation Perspective to Intangible Cultural Heritage in the Context of Digital Humanities Research”, in: Studia Ethnologica Croatica 35, 1: 115–148. DOI: 10.17234/SEC.35.8.