DH 2026

Daejeon, July 27–31

Thu, July 3016:30–18:00S107204-205
Short Paper

FAIR Memories: A Workflow for the Preservation and Collaborative Reuse of Analog Oral History Collections

Laurent Fintoni
University of Bologna, Italy · laurent.fintoni2@unibo.it
Nike del Quercio
University of Bologna, Italy · nike.delquercio2@unibo.it
Costanza Paolillo
University of Bologna, Italy · costanza.paollilo@unibo.it

While oral sources are recognized as critical scholarly and cultural assets (‘Archiving Oral History’ 2019), many collections

We should note here the inconsistency of labels used when referring to oral history—archive, corpus, and collection but also oral, speech, and audio which reflects its fragmentation across disciplines (Calamai / Frontini 2018). In the case of this paper, we have chosen to use oral history collection though archive or corpus could also be applicable.

produced before the widespread adoption of digital standards remain fragile, difficult to access, and ethically complex (Bonomo et al. 2016). Furthermore, the digitization and reuse of analogue oral history collections remains a persistent challenge (Rogers / Mcallister 2014) one that sits at the intersection of important current themes such as sustainability, reuse, and ethics (Calamai 2025). These materials, often recorded on frail analog media and governed by consent frameworks that predate contemporary privacy regulations, have high research value but frequently fall outside current Open Science standards without extensive, additional transformative work. Within the context of the humanities, the preservation and reuse of oral history collections cuts across many fields and practices and is particularly relevant to the digital humanities where a multidisciplinary approach is well suited to bringing together the different technological, curatorial, and organizational skills needed to manage and valorize these collections (Calamai / Frontini 2018). Aligning the digital transformation of oral history collections with the FAIR Data Principles (Wilkinson et al. 2016), while not straightforward, can make them more accessible both to researchers but also to the wider public, the latter a key element in aligning oral history practices with the interests and desires of the communities that contribute to them. This paper presents FAIR Memories, a Horizon Europe-funded project developing a reusable workflow for the FAIR-ification of legacy oral history collections, allowing for their transition to current Open Science standards. Using a historically important multimodal collection about Italian popular culture, the workflow is being developed jointly by domain experts and digital humanists with particular attention to the following: preservation of the recorded voice as a core epistemic dimension; standards-based interoperability; semantic enrichment; and compliance with contemporary privacy regulations. Here we present the initial stages of the project, detailing the collection used, the proposed steps of the workflow, and some of our key theoretical approaches.

To achieve its goals, FAIR Memories makes use of a historically significant but underused collection of oral history interviews on Italian popular culture, recorded in the early 1990s for the project Cultural Industries, Governments and the Public in Italy, 1938–1954 (Forgacs / Gundle 2010), sponsored by the UK’s Economic and Social Research Council. The collection consists of 117 interviews with individuals born between the 1910s and 1930s, conducted across the country in 1991 and 1992, that examine relations between cultural production, consumption, and political power before and after World War II. It remains to this day the most extensive and substantially funded oral history research on mass culture in those years ever conducted in the country, the results of which formed the basis for the book Mass Culture and Italian Society from Fascism to the Cold War (Forgacs / Gundle 2007). These interviews were preserved on audio cassettes, typed transcripts printed on paper, and photocopies of the transcripts, with various duplicates of both cassettes and paper transcripts spread between the different researchers. Textual components (over 3,000 pages of photocopied transcripts) were scanned and deposited with the UK Data Service (UKDS) in the late 2000s, alongside interview metadata (biographic and geographic data and interview dates), however the promise of the addition to this deposit of the audio material (around 90 cassettes) was never completed. The collection is currently available as safeguarded data, requiring registration with the UKDS, a process made easier for academic users but not the general public

See https://ukdataservice.ac.uk/find-data/access-conditions/

. Thus, this multimodal collection exists in a fragmented digital, and physical, state. Textual records are technically preserved, but are not easily accessible to non-academics and are composed of low-quality digital copies of physical copies of the original typed transcripts, well below accepted standards for archival records. The original copies meanwhile appear to be currently lost, with only one set of photocopies remaining available to us. Meanwhile, the recorded voices of everyday people remain tied to cassette tapes, which are copies of the originals (also currently appearing to be lost) and do not include the complete set of interviews. With only the visual counterpart of these aural objects available, the prosodic and embodied features (intonation, hesitation, pacing, involuntary speech acts) that constitute the most representative and specific part of oral testimony (Pessanha 2022, Viswanath et al. 2024) remain obscured and currently inaccessible in their written form (Portelli 2016).

As with many similar collections found across the Global North, and in particular Europe, the fragmented state of our collection reflects the inherent politics of digitization projects (Zaagsma 2023) and highlights persistent challenges: the technological obsolescence of recording media; the uneven remediation of analog materials; the difficulty in applying broad ethical and legal guidelines to oral history collections, which are characteristically defined by their uniqueness; and the gap between formal data deposit and meaningful public or scholarly access. Tackling these challenges requires collaborative, multidisciplinary work to go beyond merely creating digital surrogates to instead address how interoperability, contextualization, and reuse can benefit both scholars (for example by growing a shared body of knowledge on how to navigate ethical and legal issues) as well as the communities that contributed to these oral history collections, so as not to further deprive them of potentially meaningful resources (Garellek et al. 2020) while underscoring the importance of the process of community restitution as equally valuable to recovery and reuse (Ardolino et al. 2025). It’s also important to remember that metadata, not just digitization, is also a major point of entry to digital cultural heritage for many (Zaagsma 2023) and that in non-European contexts orality is understood “as knowledge, living and evolving” (Riaño-Alcalá 2025, 495) highlighting the need for workflows that balance preservation with access to both the oral and textual components. In this sense, our collection functions not merely as necessary data with which to test and develop FAIR-oriented remediation strategies advocated by the workflow, but as a paradigmatic example of the limits of early digital transitions in the humanities (Zaagsma 2024).

The digitization and FAIR-ification workflow we are developing is conceived, and will be delivered, within the context of the European Open Science Cloud (EOSC). As such it is thought of as a form of open method, a type of research product that is more fluid than traditional ones and is integral to the work of the Social Sciences and Humanities cluster within EOSC via its Social Sciences and Humanities Open Marketplace (SSHOMP). In the SSHOMP, workflows can be thought of as concrete examples of best research practices within a given community, focusing on accessible and reproducible processes rather than products. Our goal is to deliver our workflow in two formats: as a narrative workflow, as defined by Barbot et al. (2024) in the context of the SSHOMP, where it can be integrated with other Open Science tools and resources already included in the marketplace, adding further context to the service; and also as an interactive workflow featuring conditional logics and a web-based user interface that can more closely mirror real-life situations and allow users outside of the EOSC to evaluate their own situation and better understand how to adapt the suggested actions, tools, and reflections to their specific needs and situations.

Our proposed workflow will involve the following principal stages: (a) assessment of the existing analogue materials and their digitization (focused on text and audio), including recommendations for documenting the history of the collection and possible revision and alignment of pre-existing digitized assets to current norms; (b) normalization of textual materials using Text Encoding Initiative guidelines to enhance machine readability while preserving features specific to oral discourse, in our case by using the adapted version of the Jefferson system developed by the KiParla project

https://kiparla.it/

in order to mitigate how transcription can flatten the cultural richness of oral testimony, filtering out layers of meaning embedded in accent, register, and language switching (Alfonzetti 2009); (c) identification and anonymization of potentially sensitive content using OpenAIRE’s Amnesia

https://amnesia.openaire.eu/

, or similar tools, as an example of how to adapt current solutions to oral data collected under pre-digital ethical frameworks specifically in the European context where GDPR policies can often act as a barrier to public access; (d) data modelling options that reuse established international standards, including Dublin Core, CLARIN’s Component Metadata Infrastructure, and the Europeana Data Model, with particular attention to existing oral history profiles (Calamai et al. 2022), to facilitate interoperability with the EOSC as well as demonstrate how to align oral history collections with the metadata needs of current infrastructures; (e) semantic enrichment of textual data using the INCEpTION semantic annotation platform

https://inception-project.github.io/

, from basic annotation layers (people, places, dates) to more tailored needs based on collection specifics, such as in our case the use of annotation to reflect on the sociological aspects of our collection; and (f) guidelines for the publication and reuse of the resulting data and evaluation of their FAIR-ness, bringing together existing research and recommendations to further guide the work based on needs and limitations of the users.

As an example of the usage of the workflow, we will transform part of our collection and make it available both as a FAIR dataset via existing EOSC-integrated repositories, as well as via a website aimed at the public using Oral History Metadata Synchronizer

https://www.oralhistoryonline.org/

, allowing user interaction with both audio and textual components. Here we should note that we are aware of potential ethical and legal limitations in the Italian context for the deployment of such a solution, however we are confident that the specifics of our collection may allow us to produce a public-facing interface that can provide an example for the Italian community to work through these complex issues.

In order to make visible the limits of applying standardized solutions to historically situated and ethically constrained sources, collaboration between digital humanists and domain experts is constitutive of the workflow we are proposing to help underline how the FAIR-ification process for analog oral history collections should be thought of as not just a purely technical operation. Decisions concerning digitization, metadata granularity, transcription practices, anonymization thresholds, semantic enrichment, and access conditions must be shaped jointly by technical constraints, community involvement where and when possible, and domain-specific interpretive knowledge in order to better confront the inherent fragmented nature of these collections (Calamai / Frontini 2018) and avoid simply turning the oral in oral history into an archival source (Riaño-Alcalá 2025, 494). Semantic annotation is a particularly interesting aspect of the workflow in the context of our collection, conceived not merely as a technical enhancement but as a collaborative interpretive practice. Pre- and post-war Italy was a period marked by tumultuous, profound social and economic transformations. As such part of our annotation work will focus on how language use, including the interaction between widespread regional dialects and limited Italian-language proficiency reshaped by internal migration, and references to everyday practices of consumption with clear class connotations were expressions of belonging or exclusion from social groups (Alfonzetti 2009; Forgacs 2014; Wright 2000). In doing so we hope to show how this stage can function as a mediating layer between computational accessibility and cultural interpretation, enabling the re-emergence of historically marginalized perspectives within the constraints of responsible data reuse.

Today, with growing attention to research sustainability and Open Science as well as community restitution and marginalized experiences, reactivating analog oral history collections feels both timely and necessary. By taking a multidisciplinary approach informed by the digital humanities, FAIR Memories aims to support existing national and international efforts to enhance the discoverability and reuse of such collections while addressing the cultural and social complexity of this form of testimony and ensuring critically informed scholarly and public engagement.

Acknowledgements

The authors acknowledge the OSCARS project, which has received funding from the European Commission’s Horizon Europe Research and Innovation programme under grant agreement No. 101129751.

References
  1. Alfonzetti, Giovanna (2009): “Italiano e dialetto tra generazioni”, in: Marcato, Gianna (ed): Dialetto. Usi, Funzioni e Forme. Padova: CLEUP 241-246. https://www.academia.edu/10219387/Italiano_e_dialetto_tra_generazioni_2008_in_G_Marcato_Dialetto_Usi_Funzioni_e_Forme_CLEUP.
  2. Oral History Association (2019): Archiving Oral History: Manual of Best Practices  https://oralhistory.org/archives-principles-and-best-practices-complete-manual/ [06.05.2026].
  3. Ardolino, Fabio / Brasile, Lorenza / Calamai, Silvia (2025): “Uno stress test per il Vademecum per il trattamento delle fonti orali: l’esperienza con l’Archivio orale dell’Istituto Interregionale di Studi e Ricerche della Civiltà Appenninica”, in: L’Analisi Linguistica e Letteraria 33, 3. https://www.analisilinguisticaeletteraria.eu/index.php/ojs/article/view/795
  4. Barbot, Laure / Dolinar, Maja / Gray, Edward J. / Grisot, Cristina / Illmayer, Klaus / Kurzmeier, Michael / McGillivray, Barbara (2024): “Contextualizing Research Tools & Services Through Workflows in the SSH Open Marketplace”, in: Journal of Open Humanities Data, 10, 22. https://doi.org/10.5334/johd.192.
  5. Bonomo, Bruno / Casellato, Alessandro / Garruccio, Roberta (2016): “Maneggiare con cura: Un rapporto sulla redazione delle Buone pratiche per la storia orale”, in: Mestiere di storico : rivista della Società italiana per lo studio della storia contemporanea : 8, 2, 2016. https://digital.casalini.it/10.23744/1324.
  6. Calamai, Silvia / Frontini, Francesca (2018): “FAIR data principles and their application to speech and oral archives”, in: Journal of New Music Research 47, 4: 339–354. https://doi.org/10.1080/09298215.2018.1473449
  7. Calamai, Silvia, et al. (2022): “Ravensbrück Interviews: How to Curate Legacy Data to Make It CLARIN Compliant”, 8 July, 1–9, https://doi.org/10.3384/ecp1891.
  8. Calamai, Silvia (2025): “Why This Journal, Why Now”, in: Oral Archives Journal, 1, 5–19. https://doi.org/10.36253/oar-3334.
  9. Forgacs, David / Gundle, Stephen (2007): Mass culture and Italian society from fascism to the Cold War. Indiana University Press.
  10. Forgacs, David, / Gundle, Stephen (2010): Oral History of Cultural Consumption in Italy, 1936-1954, https://doi.org/10.5255/UKDA-SN-6479-1 [06.05.2026].
  11. Forgacs, David (2014): Italy’s Margins: Social Exclusion and Nation Formation since 1861. Cambridge University Press. https://doi.org/10.1017/CBO9781107280441
  12. Garellek, Mark / Simpson, Adrien / Roettger, Timo B. / Recasens, Daniel / Niebuhr, Oliver / Mooshammer, Christine /Michaud, Alexis / Lee, Wai-Sum / Kirby, James / Gordon, Matthew / Yu, Kristine M. (2020): “Letter to the editor: Toward open data policies in phonetics: what we can gain and how we can avoid pitfalls”, in: Journal of Speech Sciences, 9, 3–16. https://doi.org/10.20396/joss.v9i00.14955
  13. Pessanha, Francisca (2022): “Non-verbal Signals in Oral History Archives”, in: Proceedings of the 2022 International Conference on Multimodal Interaction, 668–672. https://doi.org/10.1145/3536221.3557036
  14. Portelli, Alessandro (2016): ‘What Makes Oral History Di fferent’, in: The Oral History Reader, 48–58, https://doi.org/10.4324/9781315671833-5.
  15. Riaño-Alcalá, Pilar (2025): “Oralities, Voice, and Affect in Oral History Work in the Afterlives of Violence”, in: Cave, Mark / Leydesdorff, Selma (eds): Handbook of Global Oral History, Leiden: Brill 493-518.
  16. Rogers, Irene / Mcallister, Margaret (2014): Ghosts in the archives: Exploring the challenge of reusing memories. https://acquire.cqu.edu.au/articles/conference_contribution/Ghosts_in_the_archives_exploring_the_challenge_of_reusing_memories/13433291/1
  17. Viswanath, Anargh / Gref, Michael / Hassan, Teena / Schmidt, Christoph (2024): “A Multimodal, Multilabel Approach to Recognize Emotions in Oral History Interviews”, in: 12th International Conference on Affective Computing and Intelligent Interaction (ACII) https://doi.org/10.1109/ACII63134.2024.00038
  18. Wilkinson, Mark D.  / Dumontier, Michel / Aalbersberg, IJsbrand Jan / Appleton, Gabrielle / Axton, Myles / Baak, Arie / Blomberg, Niklas / Boiten, Jan-Willem / Da Silva Santos, Luiz Bonino / Bourne, Philip E. / Bouwman, Jildau /Brookes, Anthony J. / Clark, Tim / Crosas, Mercè / Dillo, Ingrid / Dumon, Olivier /Edmunds, Scott / Evelo, Chris T. / Finkers, Richard / … Mons, Barend (2016): “The FAIR Guiding Principles for scientific data management and stewardship”, in: Scientific Data 3, 1, 160018. https://doi.org/10.1038/sdata.2016.18
  19. Wright, Sue (2000): Community and Communication: The Role of Language in Nation State Building and European Integration. Bristol: Multilingual Matters & Channel View Publications. https://doi.org/10.2307/jj.33169474 
  20. Zaagsma, Gerben (2023): “Digital History and the Politics of Digitization”, in: Digital Scholarship in the Humanities 38, 2, 830–851. https://doi.org/10.1093/llc/fqac050
  21. Zaagsma, Gerben (2024): “Facing the History Machine: Toward Histories of Digital History”, in: History of Humanities 9, 2, 451–491. https://doi.org/10.1086/731827