Daejeon, July 27–31
Digital History scholarship has quickly moved from a visual turn to a multimodal turn, where audiovisual material is of highest prominence. However, working with this rich material at scale is rarely straight-forward. This short paper presents the curating and data-engineering work supporting research on Swedish newsreels. We focus on how archival and technical constraints shape what can be remembered about twentieth-century audiovisual news. Our case is Journal Digital, a collection of over 5,000 digitised newsreels and short non-fiction films (1897–1960) held by Swedish public service institutions (Snickars 2015, 2024). These films once structured how audiences saw and heard the world, but today they are locked behind a combination of low-resolution MPEG-1 surrogates, legacy catalogue records, and on-site access restrictions.
Our contribution shows how targeted preprocessing and corpus construction can re-open such an archive as a structured memory resource. While we briefly demonstrate how our pipelines can be mobilised in concrete analyses of newsreel memory and audiovisuality, the primary focus of this short paper is on the infrastructural layer: we want to highlight the crucial – but often invisible – work of curation and pre-processing of multimodal sources that make analysis possible at all. We argue that these preparatory steps are themselves forms of memory work, shaping which traces of past audiovisual news become accessible. What can be extracted, searched, and recombined from the archive is in this sense never neutral. Rather, it determines the vantage points from which both researchers and wider publics can re-encounter these newsreels, and thus which versions of audiovisual pasts become thinkable in the first place.
Related initiatives such as the CLARIAH Media Suite (Melgar-Estrada et al. 2019), I-Media-Cities (Scipione et al. 2019), and Oiva et al.’s (2024) newsreel framework focus on access, retrieval, and content analysis; by contrast, our contribution foregrounds the preprocessing decisions themselves — treating them as forms of memory work that shape what becomes researchable.
First, we build a large-scale textual layer from sound by using SweScribe (Aspenskog / Johansson 2025b), a custom Swedish ASR pipeline, to transcribe 2,553 sound films from Journal Digital, mostly newsreels with “voice-of-God” commentary, into time-aligned text, producing the Journal Digital Corpus (JDC) of roughly 2.2 million words at around 7% WER (Aspenskog et al. 2025a). In parallel, we use energy and spectral descriptors from the audio stream to detect whether a file contains usable sound and to separate these audio-bearing reels for further multimodal preprocessing, where they are paired with visual and audio object-recognition (Johansson / Malmstedt 2026).
We also recover timestamped intertitles that structured early newsreel narration: almost 50,000 intertitles (Dupré La Tour 2005; Chaume 2020). For 4,300 films, our tool stum (Johansson 2025) processes every frame, groups visually similar frames, detects likely intertitles with an EAST text detector, and runs Tesseract OCR in both original and mirrored form (Zhou et al. 2017; Smith 2007), generating per-file subtitles. Finally, we combine these streams of time-aligned data to analyse how commentary, written titles, and moving images jointly shape what is remembered as the audiovisual experience of “news” (Chion 1994; Beck 2010), while also making visible how ASR errors, OCR failures, and heuristic filters bias which parts of this audiovisual past become easiest to retrieve.
This layered approach enables, for instance, tracing the emergence or decline of specific political or social themes across three distinct modalities—the spoken commentary, the written intertitles, and the visual scene descriptions from object recognition. This allows for a granular examination of media bias, providing evidence of how the narration may diverge from or reinforce the visual and written record. The transparency of our pipeline encourages critical reflection on the biases of the ASR and OCR tools themselves, showing how algorithmic choices inevitably privilege certain parts of the archive over others. The pipeline is thus intended to do something more than simply leveraging data.
The methods we use to process, curate, and analyse such multimodal sources are themselves part of how they will be remembered. As Gitelman / Jackson (2013) argue, data are never “raw” but always already shaped by the circumstances of their collection and processing; similarly, Manovich (2001) shows how digital “transcoding” binds cultural and computational layers together, so that technical choices inevitably inflect the cultural objects they produce. By turning disparate speakertracks, intertitles, and audiovisual signals into aligned, searchable layers, we unlock new forms of research on newsreels that relocate images and sound among textual material. Making it possible to trace how different modalities inform and relate to each other. At the same time, pre-processing has to be done with scientific transparency and with careful theoretical reflection on the complex status of audiovisual documents as both historical traces and technical artefacts. The methods of processing involved in historical work always shape the possibilities of cultural memory. Thus, this paper argues for a media-specific and context-sensitive approach to audiovisual materials: by making our curatorial and computational choices explicit and open to scrutiny, we aim to enable new forms of analysis while keeping the historical, technical, and medial specificity of these documents firmly in view.