Daejeon, July 27–31
Audio description (AD), defined as the “verbal–vocal description of visual or audiovisual content for visually impaired audiences” (Hirvonen / Wiklund 2021), is a structured practice of translating visual information into language. Museum audio description, which is the focus of the present study, aims to make visual art accessible to non-sighted visitors by providing curated verbal representations of artworks. It typically follows a standardized informational sequence: first introducing the title, author, artistic movement, and physical dimensions of the artwork, and then proceeding to a detailed description organized along a specified scanning path. Such descriptions can be delivered live, pre-recorded as audio guides, or made available online.
Although museum audio description has been practiced for several decades, it remains a comparatively understudied area. Existing work has focused primarily on professional norms (Snyder 2014), interpretive strategies and pedagogical aspects (Pawłowska / Sowińska Heim 2016, 2019; Jerzakowska 2016), and challenges of verbalizing aesthetic experience (Neves 2012; Więckowski 2014). Recent corpus- and experiment-based studies suggest that AD exhibits heightened lexical density and textual complexity (Perego 2019, 2024), metaphorical richness (Luque Olmenero / Soler Gallego 2020; Spinzi 2019; Bartolini 2023), and subjectivity (Luque Olmenero / Soler Gallego 2020), making it a promising domain for investigating systematic relations between visual structure and linguistic form. Some studies propose experimental multisensory approaches, e.g., soundpainting and tactile experience, to be incorporated in AD to enhance its evocative potential (Neves 2012; Randaccio 2020).
Its theoretical framework involves the notion of sound symbolism, or phonological iconicity, which maintains that certain phonological units or their features are systematically related to meaning and perceptual experience (Akita 2013). More broadly, this work is situated within research on multimodal iconicity, which treats non-arbitrary form–meaning correspondences as a general property of language grounded in cross-modal and embodied cognition (Perniss et al. 2010; Dingemanse et al. 2016).
Drawing on established findings in sound symbolism (e.g., the kiki/bouba effect, whereby obstruent and voiceless consonants are associated with angular shapes, while sonorant and voiced sounds align with rounded forms, cf. Köhler 1929; Ramachandran / Hubbard 2011), the study examines whether measurable visual properties of abstract paintings correlate with systematic patterns in the phonological composition of their audio descriptions. We test how visual sharpness, curvature, and complexity align with distributions of phonemes and phonological features in texts. By integrating computer vision–based analysis of artworks with phoneme-level corpus methods, we model multimodal correspondences, seeking to examining how visual experience shapes the sound patterns of language in real-life use and thus contribute to the study of intersensory translation and embodied meaning-making in cultural mediation.
The study is based on a bilingual corpus of approximately 800 museum audio descriptions (ADs), both recorded and text-based, in English and Polish, collected from over 20 major international institutions. From this dataset, all available descriptions of modern abstract artworks were included (N=191). Given the relatively limited number of abstract artworks with accessible ADs, no additional filtering was required. The abstract nature of the artworks provides inherent topical balance, with a roughly even distribution between geometric and non-geometric abstraction. Each description was paired with a digital reproduction of the corresponding artwork and relevant metadata.
Visual features were extracted using computer vision methods implemented in OpenCV (Bradski 2000), including edge density, corner density, and contour curvature and variability, capturing degrees of angularity and smoothness. Visual complexity was quantified using compression-based measures (Karjus et al. 2023). These metrics were standardized into two continuous indices: (i) a visual kiki–bouba score (sharpness vs. roundness) and (ii) visual complexity.
Phonological representations were derived using a hybrid pipeline. Recorded ADs were processed with the Montreal Forced Aligner (MFA; McAuliffe et al. 2017) to obtain phonemic alignments, while text-only descriptions were converted using MFA grapheme-to-phoneme (G2P) models, ensuring a consistent phoneme inventory. Although automated transcription may introduce minor inaccuracies (e.g. proper names), these are assumed to be non-systematic and mitigated by corpus scale and the use of aggregated feature measures. Phonemes were annotated for articulatory features based on standardized MFA mappings. Phonological measures were computed as normalized feature proportions, with text length (in phonemes) included as a covariate to control for verbosity.
Statistical analysis models visual properties as continuous outcomes predicted by phonological features. Linear mixed-effects models were fitted with visual kiki–bouba and complexity scores as dependent variables and phonological feature proportions as fixed effects. Language and museum were included as random intercepts to account for cross-linguistic and institution-specific stylistic variation. This approach identifies the phonological predictors most strongly associated with visual sharpness, smoothness, and complexity while maintaining interpretability.
Our exploratory analyses indicate systematic correspondences between visual properties of abstract paintings and phonological patterns in their audio descriptions. Higher ‘kiki’ scores, i.e., in paintings featuring sharper contours align with a higher proportion of obstruents, voiceless consonants, and front vowels in the descriptions, whereas paintings with smoother contours and lower edge density align with more frequent sonorants, voiced consonants, and back vowels. Paintings displaying higher compression-based complexity correlate with greater phonological density, more segmented language in the form of higher consonant-to-vowel ratios, and more obstruent consonants in the ADs. While the exact magnitude of these effects remains to be established, practically meaningful correspondences are defines as those that are consistent across languages and institutions and robust under control for verbosity and stylistic variation. As for AD practice, such patterns would be vital if they point to systematic tendencies in how describers linguistically encode visual properties, thereby offering a basis for making descriptive strategies more perceptually aligned.
The study contributes a computational framework for modeling multimodal iconicity in museum audio description, extending sound–shape research to real-world accessibility contexts. It positions AD as a valuable resource for studying embodied and intersensory meaning-making in cultural mediation, with implications for linguistics, digital humanities, and accessibility research.
This research was funded by the National Science Centre, Poland, under project number 2024/53/B/HS2/04292 “Reverberations: Museum Audio Description and Intersemiotic Translation Practices”.