Daejeon, July 27–31
Archives and memory sources are never neutral (Hedstrom 2002). They are shaped by power and unequal capacities to record and transmit experience (Pavone 2009), especially in contexts of displacement and political conflict (Bastian 2003). Beyond Archives is a project at the intersection of Archival Science and Digital Humanities that develops a relational digital library for sources on the post-war Julian–Dalmatian exodus to Naples. It brings institutional records and oral histories into a computable environment designed to make contested memories visible, searchable, and traversable for heterogeneous publics.
The project is guided by questions concerning the digitization of plural memory sources, the relation between institutional narratives and personal recollections, and the role of Archival Science and Digital Humanities in making this complexity legible and communicable. In this sense, it contributes to the design of archival infrastructures that connect records, communities, and interpretive practices.
Work at the intersection of Memory Studies and Archival Theory has shown that archives belong to wider memorial networks and that records, provenance, value, and representation are shaped by institutional histories and political choices (Giuva et al. 2007; Hedstrom 2010; Lian 2017; Di Marcantonio 2022). This condition is especially visible in contexts marked by displacement and political conflict, where traveling and multidirectional approaches to memory foreground movement, negotiation, and the coexistence of heterogeneous perspectives (Erll 2011; De Cesari / Rigney 2014; Rothberg 2009). In this context, institutional documentation enters a broader documentary field and requires an infrastructure able to connect traces across subjects, places, and narratives.
Archival scholarship has addressed this challenge by questioning the notion of a single monolithic “Archive” and by opening archival thought to plural documentary traditions and distributed forms of recordkeeping through the records continuum and the archival multiverse (Caswell 2016; McKemmish 2017; McKemmish et al. 2009; Gilliland 2015; Gilliland 2017). This framework supports the relation of institutional and personal traces within the same relational environment while preserving their distinct documentary logics.
Computational approaches in Digital Humanities extend this framework through formalization, requiring sources to be expressed through explicit structures and models so that they can become computable objects (Tomasi / Buzzetti 2012; Tomasi 2022). From this perspective, Knowledge Organization and archival modelling are especially relevant because they provide ways of structuring entities, relations, description, access, and reuse within a shared system (Hjørland 2008; Gnoli et al. 2006; Michetti 2015; Michetti 2023). They inform the construction of a relational library organized through entities, relations, and access points that make the corpus legible, navigable, and interoperable.
Beyond Archives addresses contested memory through the design of an infrastructure that relates different documentary perspectives while preserving their forms and contexts of production. Developed on a small-scale corpus, it uses this environment to test and document a workflow that supports comparison, navigation, and reuse within the case study and provides a methodological basis for adaptation to similar contexts of contested memory.
The first of these two connected streams consists of institutional written records from the Prefecture of Naples, drawn from six archival boxes dated 1946–1958 and currently explored through an initial set of 178 typewritten administrative records from the first box. The second consists of a small set of semi-structured oral interviews collected in collaboration with community associations and governed by consent and reuse agreements. Together, these materials form a memorial ecosystem in which distinct documentary perspectives become traversable through a relational layer of entities and access points, including persons, places, institutions, and accommodation sites. This structure supports movement across the corpus while preserving differences of provenance, form, and descriptive logic.
That relational layer is formalized through a TEI-based modeling environment conceived as a negotiated space. TEI-XML provides the formal layer of the project, supporting the representation of written and oral sources within a common relational framework while preserving their documentary structures, descriptive logics, and conditions of use (Tomasi / Buzzetti 2012; Tomasi 2022). TEI headers record provenance, production context, repository links, and access or reuse conditions, while the encoded texts share a common layer of entities and relations. TEIPublisher exposes this formalization as relational navigation, allowing users to move among facsimiles, entities, and testimonies through a common set of access points.
The workflow is articulated in three stages: digitization, formalization, and publication. For the written stream, formalization depends on a coupled recognition layer that combines semantic layout segmentation and Automatic Text Recognition. It focuses on mid-twentieth-century typewritten Italian administrative records whose apparent regularity coexists with stamps, annotations, and material degradation. These features require corpus-specific refinement of public models to support text extraction and document-aware encoding (Chiffoleau 2025; Ströbel et al. 2022). In eScriptorium, segmentation identifies semantically relevant regions and lines, and ATR is then performed on that structured page, beginning from CATMuS-Print and iteratively fine-tuning it on the Italian corpus (Gabay / Clérice 2024). The resulting PAGE-XML and ALTO-XML outputs preserve layout, coordinates, and semantic segmentation, and guide the subsequent TEI encoding workflow. The written corpus thus becomes structured and reusable for downstream modelling and publication. Evaluation combines quantitative metrics with qualitative assessment and is framed in terms of fitness-for-purpose (Bubula et al. 2025).
FAIR and open science principles inform the project throughout, shaping the choice of open tools, the documentation of the workflow, and the planned release of models, paradata, and selected ground-truth (Wilkinson et al. 2016). Access and reuse are calibrated according to the nature and provenance of the sources: oral materials are governed by consent and reuse agreements, while the release of written records is defined in collaboration with the State Archives of Naples. In this way, workflow transparency and documentation remain integrated with source-aware access conditions.
By treating archival and oral traces as computable objects within a documented workflow, Beyond Archives develops a relational library for contested memory grounded in archival rigor, formal modelling, and calibrated access conditions. Its small-scale design supports close methodological control and sustained documentation of the pipeline, including models, paradata, and structured outputs for future reuse. This combination offers an adaptable framework for relating documentary plurality and supporting navigation and reuse in similar corpora shaped by contested memory.