Daejeon, July 27–31
We should note here the inconsistency of labels used when referring to oral history—archive, corpus, and collection but also oral, speech, and audio which reflects its fragmentation across disciplines (Calamai / Frontini 2018). In the case of this paper, we have chosen to use oral history collection though archive or corpus could also be applicable.
produced before the widespread adoption of digital standards remain fragile, difficult to access, and ethically complex (Bonomo et al. 2016). Furthermore, the digitization and reuse of analogue oral history collections remains a persistent challenge (Rogers / Mcallister 2014) one that sits at the intersection of important current themes such as sustainability, reuse, and ethics (Calamai 2025). These materials, often recorded on frail analog media and governed by consent frameworks that predate contemporary privacy regulations, have high research value but frequently fall outside current Open Science standards without extensive, additional transformative work. Within the context of the humanities, the preservation and reuse of oral history collections cuts across many fields and practices and is particularly relevant to the digital humanities where a multidisciplinary approach is well suited to bringing together the different technological, curatorial, and organizational skills needed to manage and valorize these collections (Calamai / Frontini 2018). Aligning the digital transformation of oral history collections with the FAIR Data Principles (Wilkinson et al. 2016), while not straightforward, can make them more accessible both to researchers but also to the wider public, the latter a key element in aligning oral history practices with the interests and desires of the communities that contribute to them. This paper presents FAIR Memories, a Horizon Europe-funded project developing a reusable workflow for the FAIR-ification of legacy oral history collections, allowing for their transition to current Open Science standards. Using a historically important multimodal collection about Italian popular culture, the workflow is being developed jointly by domain experts and digital humanists with particular attention to the following: preservation of the recorded voice as a core epistemic dimension; standards-based interoperability; semantic enrichment; and compliance with contemporary privacy regulations. Here we present the initial stages of the project, detailing the collection used, the proposed steps of the workflow, and some of our key theoretical approaches.See https://ukdataservice.ac.uk/find-data/access-conditions/
. Thus, this multimodal collection exists in a fragmented digital, and physical, state. Textual records are technically preserved, but are not easily accessible to non-academics and are composed of low-quality digital copies of physical copies of the original typed transcripts, well below accepted standards for archival records. The original copies meanwhile appear to be currently lost, with only one set of photocopies remaining available to us. Meanwhile, the recorded voices of everyday people remain tied to cassette tapes, which are copies of the originals (also currently appearing to be lost) and do not include the complete set of interviews. With only the visual counterpart of these aural objects available, the prosodic and embodied features (intonation, hesitation, pacing, involuntary speech acts) that constitute the most representative and specific part of oral testimony (Pessanha 2022, Viswanath et al. 2024) remain obscured and currently inaccessible in their written form (Portelli 2016).https://kiparla.it/
in order to mitigate how transcription can flatten the cultural richness of oral testimony, filtering out layers of meaning embedded in accent, register, and language switching (Alfonzetti 2009); (c) identification and anonymization of potentially sensitive content using OpenAIRE’s Amnesiahttps://amnesia.openaire.eu/
, or similar tools, as an example of how to adapt current solutions to oral data collected under pre-digital ethical frameworks specifically in the European context where GDPR policies can often act as a barrier to public access; (d) data modelling options that reuse established international standards, including Dublin Core, CLARIN’s Component Metadata Infrastructure, and the Europeana Data Model, with particular attention to existing oral history profiles (Calamai et al. 2022), to facilitate interoperability with the EOSC as well as demonstrate how to align oral history collections with the metadata needs of current infrastructures; (e) semantic enrichment of textual data using the INCEpTION semantic annotation platformhttps://inception-project.github.io/
, from basic annotation layers (people, places, dates) to more tailored needs based on collection specifics, such as in our case the use of annotation to reflect on the sociological aspects of our collection; and (f) guidelines for the publication and reuse of the resulting data and evaluation of their FAIR-ness, bringing together existing research and recommendations to further guide the work based on needs and limitations of the users.https://www.oralhistoryonline.org/
, allowing user interaction with both audio and textual components. Here we should note that we are aware of potential ethical and legal limitations in the Italian context for the deployment of such a solution, however we are confident that the specifics of our collection may allow us to produce a public-facing interface that can provide an example for the Italian community to work through these complex issues.