DH 2026

Daejeon, July 27–31

Fri, July 3109:00–10:30S038101-102
Long Paper

Contested Memories as Data: Digital Humanities and the Spanish Civil War

Gustavo Candela
University of Alicante, Spain · gcandela@ua.es
Paul Joseph Spence
King's College London, United Kingdom · paul.spence@kcl.ac.uk

Introduction

The Spanish Civil War (1936–1939) was one of the most significant historic events of the twentieth century, widely viewed as a battleground between fascism, communism and liberal democracy which formed a prelude to the Second World War. One of the challenges in documenting a historic event such as the Spanish Civil War (SCW), site of highly contested and geographically dispersed memories, is the very nature of the fragmentation in terms of geographies, memories and evidence. In the words of Cazorla et al., “the Public History of the Spanish Civil War and its aftermath is characterised by the physical and institutional isolation of the entities that exist”, making it a field well-suited for efforts to connect and integrate disparate sources and in many cases published as data silos. However, much of the data about the conflict and its legacy is not readily available in digital form and there is a wide variety of formal and informal documentation in different formats, with no central or federated system for searching or connecting across the various systems holding information about the conflict.

One common approach to connecting dispersed and heterogeneous historical archives is to use Linked Open Data, an approach which renders data interoperable across different platforms and research fields (Berners-Lee et al., 2001). Previous SCW work has tackled this topic covering different angles and perspectives, and has aimed to capture textual and audiovisual testimonies of victims, activists, survivors, and witnesses of the Spanish Civil War and its aftermath (Bocanegra-Barbecho, 2020; Martin-Cabrera, 2019). However, these works are often limited in their adoption of FAIR principles (Wilkinson, 2016), hindering the extension and integration of this research data with new initiatives.

Existing datasets covering the SCW as a main topic include archives such as Archivo de la Democracia at the University of Alicante

https://archivodemocracia.ua.es/

and Portal de Archivos Españoles.

https://pares.cultura.gob.es/inicio.html

Other repositories using advanced metadata technologies based on LOD also cover the SCW such Biblioteca Virtual Miguel de Cervantes and the national libraries of Spain and France.

Over the last decade Cultural Heritage (CH) organisations have explored new ways to make their digital collections available to encourage their reuse in innovative and creative ways (Mahey et al., 2019; Padilla et al., 2023; Carroll et al., 2020). The objective of our study was to examine how cross-collection approaches based on these principles may be applied to historical research, in particular in relation to contested memories, using an experimental framework to extract datasets suitable for computational use from data structured as knowledge graphs. The purpose of the framework is to facilitate wider re-use and analysis of existing datasets in historical research and to help create the conditions for closer collaboration between cultural heritage, memory studies and historical research agendas in this area in future.

There has however been relatively little work on analysing digital methods, data curation or research infrastructure, or concerning the publication of digital collections suitable for computational use related to the Spanish Civil War and post-war Exile. The purpose of this work is to provide a reproducible framework to leverage the use of KGs to extract Collections as Data. The main contributions of this work are thus as follows: i) a reproducible framework to extract Collections as Data according to specific themes using existing KGs; ii) an analysis of available datasets focused on the SCW; and iii) the reproducible code and examples provided, which can be applied to SCW research, or indeed wider historical or memory contestation research. These contributions are intended to encourage CH institutions and historical or DH researchers to publish and reuse their data as Collections as Data, as well as to explore some of the conditions for closer collaborations between stakeholders in future.

Methodology

We follow a ‘Collections as Data’-based methodology, leveraging the use of existing knowledge graphs to extract research data in three steps: i) identification of sources; ii) data extraction and modelling; and iii) publication and reuse. The steps are described below.

The first step corresponds to the identification of sources according to the topic. We presented earlier a selection of projects and digital collections that can be analysed. They cover a wide diversity of resources including relevant people, bibliographic information and metadata, images, ships, military camps and general information related to a particular historical domain (the Spanish Civil War). Most of the examples analysed here consist of basic HTML-based websites containing only detailed human-readable information, in other words, it is relatively unstructured and less machine actionable and FAIR than data in more standards-based systems, as is characteristic of much data in this field.

In the next step we perform data extraction and modelling using the original sources identified in the previous step. Controlled vocabularies and ontologies such as Schema.org and CIDOC-CRM can be employed to describe metadata and datasets with semantic annotation which helps to facilitate connections between repositories. In addition, known repositories such as Wikidata and Wikibase enable the publication of FAIR data as well as the enrichment with external repositories (Candela et al., 2024; Lindemann et al., 2025).

The last step corresponds to the publication and reuse of the extracted datasets for future reuse. Several initiatives can be employed to increase the visibility of the datasets such as Github -based documentation, code and data, curated metadata (Alkemade et al., 2025) on the Social Sciences and Humanities Open Marketplace, dedicated platforms to enable data download data as Zip files or research data initiatives in CH institutions such as the British Library Research Repository. In this sense, detailed workflows have been recently released to publish machine-actionable collections (Candela et al, 2025).

Results

This section describes the results obtained after applying the framework proposed in this work. Three research scenarios were defined according to the following topics: artists migrating to foreign countries; posters of the SCW;

https://pares.cultura.gob.es/cartelesGC/

and the International Train Station in Canfranc during the Second World War. These examples cover a wide diversity of content in terms of available data, size, format and reuse examples.

A collection of Jupyter Notebooks was created in order to provide reproducible code and documentation to show how to build, enrich and extract the data from known sources such as the ones described earlier. The code can be run in cloud services such as Binder and European Open Science Cloud dedicated service for Notebooks.

https://open-science-cloud.ec.europa.eu/services/interactive-notebooks

Discussion

Data-driven approaches to curating and researching historical conflicts and memories offer significant potential, but we also outline the challenges for their development and interpretation. Data curation must overcome archival silences or asymmetries, but how to negotiate questions of provenance, bias and perspective? How to capture digitisation bias and uneven coverage, or the contrast between personal testimony and official archives? How to integrate both researcher and community-led perspectives in an ethics-driven manner? And how to avoid collapsing epistemic and political differences? How can existing cloud services such as Wikibase be employed to describe SCW content and to facilitate FAIR data? Are institutions ready to start using new data research infrastructures? In this paper, we ground the framework used and the experimental visualizations produced as a result in analysis of these questions and reconceptualise Collections as Data as a critical memory practice that can expose and negotiate epistemic, curatorial and ethical tensions inherent in researching contested historical memory, using the Spanish Civil War as a case study.

Conclusions

We followed in this work a reproducible framework for the reuse of data relating to contested historical memory which, borrowing in particular from cultural heritage data practices, explores the use of knowledge graphs according to Collections as Data principles. The framework was applied to a selection of research scenarios focused on the Spanish Civil War topic. Results show that there is still room for improvement concerning institutions and the provision of machine-actionable collections. Future work to be explored includes the exploration of new data research infrastructures such as the European data space for cultural heritage and cloud services to run reproducible code requiring high-performance resources.

Reproducible code and examples

https://github.com/hibernator11/Spanish-Civil-War-KGs

https://github.com/hibernator11/wikidata-queries

References
  1. Alkemade, H., Candela, G., Claeyssens, S., Colavizza, G., Eren, S., Freire, N., Irollo, A., Isaac, A., Lehmann, J., Neudecker, C., Osti, G., van Strien, D., & Wevers, M. (2025). Datasheets for Digital Cultural Heritage Datasets (Versión 2). Zenodo.https://doi.org/10.5281/zenodo.15828222
  2. Berners-Lee, T., Hendler, J. and Lassila, O., The Semantic Web, Scientific American Magazine 284 (2001).
  3. Bocanegra-Barbecho, L.. (2020). Ten years recovering the memory of Republican Exile with citizen collaboration. The results of e-xiliad@s project: a perspective from digital humanities and digital public history. https://doi.org/10.17613/6ppd-qq35
  4. Candela, G., Chambers, S., Irollo, A., Freire, N., Dritsou, V., Isaac, A., Benardou, A., Garnett, V., & Tasovac, T. (2025). “A Workflow to publish Collections as Data: looking back at Europeana.eu and forward to the common European data space for cultural heritage”, in: Transformations: A DARIAH Journal, Version 1, Vol. 965. https://doi.org/10.46298/transformations.14774
  5. Carroll, S.R., Garba, I., Figueroa-Rodríguez, O.L., Holbrook, J., Lovett, R., Materechera, S., Parsons, M., Raseroka, K., Rodriguez-Lonebear, D., Rowe, R., Sara, R., Walker, J.D., Anderson, J. and Hudson, M. (2020) “The CARE Principles for Indigenous Data Governance”, in: Data Science Journal, 19(1), p. 43.
  6. Cazorla-Sánchez, A. and Shubert, A. (2018). Sites Without Memory and Memory Without Sites: On the Failure of the Public History of the Spanish Civil War. In Public Humanities and the Spanish Civil War: Connected and Contested Histories, Alison Ribeiro de Menezes, Antonio Cazorla-Sánchez and Adrian Shubert (eds.). Springer International Publishing, Cham, 19–43. https://doi.org/10.1007/978-3-319-97274-9_2
  7. Candela, G., Cuper, M., Holownia, O., Gabriëls, N., Dobreva, M. and Mahey, M. 2024. “A Systematic Review of Wikidata in GLAM Institutions: a Labs Approach”, in: Linking Theory and Practice of Digital Libraries: 28th International Conference on Theory and Practice of Digital Libraries, TPDL 2024, Ljubljana, Slovenia, September 24–27, 2024, Proceedings, Part II. Springer-Verlag, Berlin, Heidelberg, 34–50. https://doi.org/10.1007/978-3-031-72440-4_4
  8. Lindemann, D., Candela, G., Pellizzari di San Girolamo, C. C., Olea, I., Varvantakis, C., Schöch, C., Santiago Faria, A., Moitinho de Almeida, V., Assis, T., & Marchetti, A. (2025, June 18). MediaWiki-based tools and services in Digital Humanities workflows. DARIAH Annual Event 2025 (DARIAH-AE2025), Göttingen, Germany. Zenodo.https://doi.org/10.5281/zenodo.15690771
  9. Mahey, M., Al-Abdulla, A., Ames, S., Bray, P., Candela, G., Chambers, S., Derven, C., Dobreva-McPherson, M., Gasser, K., Karner, S., Kokegei, K., Laursen, D., Potter, A., Straube, A., Wagner, S-C. and Wilms, L., with forewords by: Al-Emadi, T. A., Broady-Preston, J., Landry, P. and Papaioannou, G. (2019) Open a GLAM Lab. Digital Cultural Heritage Innovation Labs, Book Sprint, Doha, Qatar, 23-27 September, 2019
  10. Martin-Cabrera, L. and Davis, A. (2019). “The Spanish Civil War Memory project: Constructing and enhancing a digital archive”, in: Bulletin for Spanish and Portuguese historical studies 43, 1. https://asphs.net/article/the-spanish-civil-war-memory-project-constructing-and-enhancing-a-digital-archive/
  11. Padilla, T., Scates Kettler, H., Varner, S., & Shorish, Y. (2023). Vancouver Statement on Collections as Data. Zenodo.https://doi.org/10.5281/zenodo.8342171
  12. Wilkinson, M., Dumontier, M., Aalbersberg, I. et al. “The FAIR Guiding Principles for scientific data management and stewardship”, in: Sci Data 3, 160018 (2016). https://doi.org/10.1038/sdata.2016.18