DH 2026

Daejeon, July 27–31

Thu, July 3011:00–12:30S079107
Short Paper

Speaking Shapes: Modeling Sound–Shape Iconicity in Museum Audio Description

Szymon Pindur
Jagiellonian University in Krakow, Poland · szymon.pindur@doctoral.uj.edu.pl
Agata Hołobut
Jagiellonian University in Krakow, Poland · agata.holobut@uj.edu.pl

Introduction

Audio description (AD), defined as the “verbal–vocal description of visual or audiovisual content for visually impaired audiences” (Hirvonen / Wiklund 2021), is a structured practice of translating visual information into language. Museum audio description, which is the focus of the present study, aims to make visual art accessible to non-sighted visitors by providing curated verbal representations of artworks. It typically follows a standardized informational sequence: first introducing the title, author, artistic movement, and physical dimensions of the artwork, and then proceeding to a detailed description organized along a specified scanning path. Such descriptions can be delivered live, pre-recorded as audio guides, or made available online.

Although museum audio description has been practiced for several decades, it remains a comparatively understudied area. Existing work has focused primarily on professional norms (Snyder 2014), interpretive strategies and pedagogical aspects (Pawłowska / Sowińska Heim 2016, 2019; Jerzakowska 2016), and challenges of verbalizing aesthetic experience (Neves 2012; Więckowski 2014). Recent corpus- and experiment-based studies suggest that AD exhibits heightened lexical density and textual complexity (Perego 2019, 2024), metaphorical richness (Luque Olmenero / Soler Gallego 2020; Spinzi 2019; Bartolini 2023), and subjectivity (Luque Olmenero / Soler Gallego 2020), making it a promising domain for investigating systematic relations between visual structure and linguistic form. Some studies propose experimental multisensory approaches, e.g., soundpainting and tactile experience, to be incorporated in AD to enhance its evocative potential (Neves 2012; Randaccio 2020).

Research aims

Its theoretical framework involves the notion of sound symbolism, or phonological iconicity, which maintains that certain phonological units or their features are systematically related to meaning and perceptual experience (Akita 2013). More broadly, this work is situated within research on multimodal iconicity, which treats non-arbitrary form–meaning correspondences as a general property of language grounded in cross-modal and embodied cognition (Perniss et al. 2010; Dingemanse et al. 2016).

Drawing on established findings in sound symbolism (e.g., the kiki/bouba effect, whereby obstruent and voiceless consonants are associated with angular shapes, while sonorant and voiced sounds align with rounded forms, cf. Köhler 1929; Ramachandran / Hubbard 2011), the study examines whether measurable visual properties of abstract paintings correlate with systematic patterns in the phonological composition of their audio descriptions. We test how visual sharpness, curvature, and complexity align with distributions of phonemes and phonological features in texts. By integrating computer vision–based analysis of artworks with phoneme-level corpus methods, we model multimodal correspondences, seeking to examining how visual experience shapes the sound patterns of language in real-life use and thus contribute to the study of intersensory translation and embodied meaning-making in cultural mediation.

Methods

The study is based on a bilingual corpus of approximately 800 museum audio descriptions (ADs), both recorded and text-based, in English and Polish, collected from over 20 major international institutions. From this dataset, all available descriptions of modern abstract artworks were included (N=191). Given the relatively limited number of abstract artworks with accessible ADs, no additional filtering was required. The abstract nature of the artworks provides inherent topical balance, with a roughly even distribution between geometric and non-geometric abstraction. Each description was paired with a digital reproduction of the corresponding artwork and relevant metadata.

Visual features were extracted using computer vision methods implemented in OpenCV (Bradski 2000), including edge density, corner density, and contour curvature and variability, capturing degrees of angularity and smoothness. Visual complexity was quantified using compression-based measures (Karjus et al. 2023). These metrics were standardized into two continuous indices: (i) a visual kiki–bouba score (sharpness vs. roundness) and (ii) visual complexity.

Phonological representations were derived using a hybrid pipeline. Recorded ADs were processed with the Montreal Forced Aligner (MFA; McAuliffe et al. 2017) to obtain phonemic alignments, while text-only descriptions were converted using MFA grapheme-to-phoneme (G2P) models, ensuring a consistent phoneme inventory. Although automated transcription may introduce minor inaccuracies (e.g. proper names), these are assumed to be non-systematic and mitigated by corpus scale and the use of aggregated feature measures. Phonemes were annotated for articulatory features based on standardized MFA mappings. Phonological measures were computed as normalized feature proportions, with text length (in phonemes) included as a covariate to control for verbosity.

Statistical analysis models visual properties as continuous outcomes predicted by phonological features. Linear mixed-effects models were fitted with visual kiki–bouba and complexity scores as dependent variables and phonological feature proportions as fixed effects. Language and museum were included as random intercepts to account for cross-linguistic and institution-specific stylistic variation. This approach identifies the phonological predictors most strongly associated with visual sharpness, smoothness, and complexity while maintaining interpretability.

Results

Our exploratory analyses indicate systematic correspondences between visual properties of abstract paintings and phonological patterns in their audio descriptions. Higher ‘kiki’ scores, i.e., in paintings featuring sharper contours align with a higher proportion of obstruents, voiceless consonants, and front vowels in the descriptions, whereas paintings with smoother contours and lower edge density align with more frequent sonorants, voiced consonants, and back vowels. Paintings displaying higher compression-based complexity correlate with greater phonological density, more segmented language in the form of higher consonant-to-vowel ratios, and more obstruent consonants in the ADs. While the exact magnitude of these effects remains to be established, practically meaningful correspondences are defines as those that are consistent across languages and institutions and robust under control for verbosity and stylistic variation. As for AD practice, such patterns would be vital if they point to systematic tendencies in how describers linguistically encode visual properties, thereby offering a basis for making descriptive strategies more perceptually aligned.

Contribution

The study contributes a computational framework for modeling multimodal iconicity in museum audio description, extending sound–shape research to real-world accessibility contexts. It positions AD as a valuable resource for studying embodied and intersensory meaning-making in cultural mediation, with implications for linguistics, digital humanities, and accessibility research.

Funding

This research was funded by the National Science Centre, Poland, under project number 2024/53/B/HS2/04292 “Reverberations: Museum Audio Description and Intersemiotic Translation Practices”.

References
  1. Akita, Kimi (2013): “The lexical iconicity hierarchy and its grammatical correlates”, in: Elleström, Lars / Fischer, Olga / Ljungberg, Christina (eds.): Iconic Investigations. Amsterdam: John Benjamins: 331–350.
  2. Bartolini, Chiara (2023): “Museum Audio Descriptions vs. General Audio Guides: Describing or Interpreting Cultural Heritage?”, in: Journal of Audiovisual Translation 6: 77–98.
  3. Bradski, Gary (2000): “The OpenCV Library”, in: Dr. Dobb’s Journal of Software Tools.
  4. Dingemanse, Mark / Blasi, Damián E. / Lupyan, Gary / Christiansen, Morten H. / Monaghan, Padraic (2016): “Arbitrariness, iconicity, and systematicity in language”, in: Trends in Cognitive Sciences 20, 8: 603–615.
  5. Fryer, Louise (2013): Putting It into Words: The Impact of Visual Impairment on Perception, Experience and Presence. Ph.D. thesis, University of London https://research.gold.ac.uk/id/eprint/10152/1/PSY_thesis_Fryer_2013.pdf [05.05.2026].
  6. Hirvonen, Maija / Wiklund, Mari (2021): “From image to text to speech: The effects of speech prosody on information sequencing in audio description”, in: Text and Talk 41: 309–334.
  7. Jerzakowska, Beata (2016): Posłuchać obrazów. Podręcznik z audiodeskrypcją do reprodukcji malarskich, uzupełniający kształcenie literackie i językowe uczniów niewidomych. Łódź: Rys.
  8. Karjus, Andres / Canet Solà, Marc / Ohm, Thomas et al. (2023): “Compression ensembles quantify aesthetic complexity and the evolution of visual art”, in: EPJ Data Science 12: 21.
  9. Köhler, Wolfgang (1929): Gestalt Psychology. New York: Liveright.
  10. Luque Olmenero, María O. / Soler Gallego, Silvia (2020): “Metaphor as Creativity in Audio Descriptive Tours for Art Museums: From Description to Practice”, in: Journal of Audiovisual Translation 3: 64–78.
  11. McAuliffe, Michael / Socolof, Michaela / Mihuc, Sarah / Wagner, Michael / Sonderegger, Morgan (2017): “Montreal Forced Aligner: trainable text-speech alignment using Kaldi”, in: Proceedings of the 18th Conference of the International Speech Communication Association: 498–502.
  12. Neves, Josélia (2012): “Multi-Sensory Approaches to (Audio) Describing the Visual Arts”, in: MonTI. Monografías de Traducción e Interpretación: 277–293.
  13. Pawłowska, Anna / Sowińska-Heim, Julia (2016): Audiodeskrypcja dzieł sztuki: metody, problemy, przykłady. Łódź: Wydawnictwo Uniwersytetu Łódzkiego.
  14. Pawłowska, Anna / Sowińska-Heim, Julia (2019): Osoby z niepełnosprawnością i sztuka. Udostępnianie, percepcja, integracja. Łódź: Wydawnictwo Uniwersytetu Łódzkiego.
  15. Perego, Elisa (2019): “Into the language of museum audio descriptions: a corpus-based study”, in: Perspectives: Studies in Translation Theory and Practice 27: 333–349.
  16. Perego, Elisa (2024): Audio Description for the Arts: A Linguistic Perspective. London: Routledge.
  17. Perniss, Pamela / Thompson, Robin L. / Vigliocco, Gabriella (2010): “Iconicity as a general property of language”, in: Frontiers in Psychology 1: 227.
  18. Ramachandran, Vilayanur S. / Hubbard, Edward M. (2001): “Synaesthesia—a window into perception, thought and language”, in: Journal of Consciousness Studies 8, 12: 3–34.
  19. Snyder, Joel (2014): The Visual Made Verbal: A Comprehensive Training Manual and Guide to the History and Applications of Audio Description. Indianapolis: Dog Ear Publishing.
  20. Spinzi, Cinzia (2019): “A Cross-Cultural Study of Figurative Language in Museum Audio Descriptions: Implications for Translation”, in: Lingue e Linguaggi 33: 303–316.
  21. Więckowski, Robert (2014): “Audiodeskrypcja piękna”, in: Przekładaniec 28: 109–123.