DH 2026

Daejeon, July 27–31

Wed, July 2914:00–15:30S019105
Long Paper

Computational Phonosemantics at Scale: Measuring Sound-to-Meaning Mappings in English and Polish with Gradient Boosting and GLMs

Szymon Pindur
Jagiellonian University in Krakow, Poland · szymon.pindur@doctoral.uj.edu.pl

Introduction

The present study undertakes a detailed empirical investigation of phonosemantic systematicity — statistically identifiable regularities in the mapping between phonological form and multidimensional experiential meaning — across the general vocabulary of English and Polish. Rooted in the classic opposition between Saussurean arbitrariness (Saussure 1966/1916) and Peircean iconicity (Peirce 1931–1958), the idea of phonosemantic associations echoes the ancient physei/thesei debate on the natural versus conventional nature of the sign (Plato 1961).

The core hypothesis addressed in the research posits that vocabulary in English and Polish manifests a certain degree of systematic associations between phonological units (phonemes, articulatory features, phoneme sequences) and a wide array of perceptual, sensorimotor, affective, and cognitive semantic dimensions, with partial cross-linguistic convergence driven by universal perceptual mechanisms. This aligns with a view that integrates iconicity and arbitrariness as complementary design features of language: the former facilitating semantic clustering, word learning, and cognitive efficiency, and the latter maximizing distinctiveness among semantically close items (Lockwood / Dingemanse 2015; Dingemanse et al. 2015). Though direct sound imitation is largely absent from general vocabulary, analogical mappings — potentially rooted in articulatory gestures, acoustic properties of speech sounds (e.g., frequency, amplitude, periodicity), or perceptual entropy (Pindur 2025) — may result in some systematic patterns.

Such mappings draw on experiential features shared across modalities: high/low pitch, abrupt/continuous onset, or tonal/aperiodic quality, which can in turn extend metaphorically to higher-order dimensions like valence, potency, or activity (Sidhu / Pexman 2018). Several studies to date have addressed such phonosemantic correlations utilizing statistical methods. These include analyses of systematic sound-to-meaning mappings within single languages (e.g., Monaghan et al. 2014; Gutiérrez et al. 2016; Winter / Perlman 2021, who applied random forests to assess the simultaneous contribution of all English phonemes to perceived size), domain-specific classification tasks (e.g., Ngai et al. 2024, using XGBoost to predict gender from phonemes in Japanese names, revealing language-specific patterns such as /i/ → masculinity and /k/ → femininity), as well as cross-linguistic investigations of phonosemantic biases across thousands of languages (Blasi et al. 2016; Erben Johansson et al. 2020). Such work typically assumes that consistent cross-linguistic patterns in form-meaning associations reflect non-arbitrary, perceptually grounded regularities. While these approaches have revealed broad tendencies — particularly in size, shape, and emotional valence — they often rely on isolated semantic categories (e.g., gender classification in Ngai et al. 2024) and/or limited phonological representations, leaving large-scale sound–meaning mappings across higher-dimensional experiential semantic spaces yet to be systematically explored.

By constructing parallel phonosemantic profiles of English and Polish general lexis, our study aims to quantify the prevalence, strength, and cross-linguistic stability of these patterns, leveraging modern advances in NLP and data analysis. We thus strive to provide wider insight into the way speech sounds map onto a more diverse set of semantic dimensions in the general vocabulary of the two languages in question. English and Polish were selected as related Indo-European languages with substantially different phonological structures, enabling controlled cross-linguistic comparison of phonosemantic patterns.

Materials and Methods

The methodological framework we apply to this end integrates distributional semantic modeling utilizing transformer language models, high-dimensional phonological parsing, and predictive machine learning. The semantic structure is operationalized using 65 continuous experiential semantic dimensions as delineated by Binder et al. (2016), spanning sensory (vision, audition, touch, etc.), motor, spatial, temporal, affective, social, and cognitive domains. These dimensions are derived from human ratings (0–6 scale) for 535 high-frequency English nouns, verbs, and adjectives, reflecting neurobiologically grounded aspects of human experience of the world. For example, words such as ambulance receive high predicted values for dimensions such as Vision, Sound, and Attention.

To scale this framework to wider lexica beyond the original English seed lexicon, we have fine-tuned a BERT large model (Devlin et al. 2019) and its Polish equivalent HerBERT (Mroczkowski et al. 2021) with a custom regression head to predict the 65-dimensional Binder feature ratings. The models have been trained and cross-validated on ~20,000 example sentences split into train and validation sets containing the 535 seed words (and their Polish equivalents) together with their corresponding ratings from the original human-generated dataset. The original English Binder ratings were transferred cross-linguistically by aligning English seed words with manually selected Polish semantic equivalents, with translation choices informed by the original semantic ratings to disambiguate potentially polysemous items. The procedure assumes approximate cross-linguistic stability of the core experiential semantic space represented by culturally widespread concepts (e.g., ambulance, child, book), while culturally specific items were replaced with semantically analogous Polish equivalents to better approximate the original semantic dimensions.

The fine-tuned models achieve near-human accuracy (MSE ≈ 0.04; median R² ≈ 0.974) on unseen validation data, thus surpassing previous attempts at generalizing Binder semantic ratings to wider vocabulary (e.g., Turton et al. 2021). External validation against Lancaster Sensorimotor Norms for ~40,000 English words further demonstrated substantial correlations between predicted and human ratings across analogous sensory dimensions, confirming robust generalization beyond the original Binder dataset.

Semantic ratings are computed for large-scale general lexica (~40,000 unique words for each language). The English vocabulary is sourced from the intersection of the English Lexicon Project morphological database (Balota et al. 2007) and the CMUdict phonetic dictionary (CMUDict 2014). The Polish vocabulary is assembled from the Polish Wiktionary scrape database (Ylönen 2022) and the Morphynet Polish morphology database (Batsuren et al. 2021).

For each target word, the model processes up to about 100 sentences sourced from the 1B Word Benchmark (Chelba et al. 2013) for English and from the Pol Eval Language Modeling Task corpus for Polish (Wojdyga 2018), generating context-sensitive 65-dimensional vectors. Ratings are POS-disambiguated and averaged to produce stable semantic profiles per word form. This yields a high-resolution, experientially grounded semantic space for downstream phonosemantic modeling.

Phonological and morphological variables are obtained through transforming standardized IPA transcriptions (sourced from the CMUDict for English and Wiktionary for Polish) into sparse binary feature matrices encoding: (i) language-specific phoneme inventories, (ii) universal articulatory features (place, manner, voicing, vowel height/backness), (iii) phoneme bigrams, (iv) derivational and inflectional morphemes, and (v) coarse-grained POS categories.

In predictive modeling, we utilize XGBoost regression with separate feature sets for each phonology group (phonemes, features, bigrams) to identify the strongest predictors for each of the 65 semantic dimensions. Hyperparameters are optimized through randomized search, with early stopping on held-out validation data. Permutation feature importance identifies the most contributory predictors per dimension. To reduce confounding effects of shared morphology, phonological predictors exhibiting substantial collinearity with morphological variables are excluded from subsequent statistical analyses. This procedure helps distinguish potentially iconic effects from correlations arising through shared morphology or lexical clustering.

Generalized Linear Models — adaptively selected for the underlying data distributions — are fitted using top-selected predictors. The statistically significant phonological predictors are identified for each semantic dimension employing the Bonferroni correction for multiple testing. The effect sizes across all features are calculated using Cliff’s Delta (Meissel / Yao 2024), which provides a robust, non-parametric effect size estimation method.

Results

The results reveal several robust and consistent phonosemantic systematicity effects. The key patterns include, for example, high front vowels (/ɪ/, /iː/) negatively correlating with Vision (words rich in /ɪ/ evoke lower visual salience), nasals (/m/, /n/) positively linked to Duration and Weight, and sibilants (/s/, /z/) correlating with higher levels of Audition and Unpleasant. Comparable effects also emerge in Polish, for example with the cluster /zb-/ showing elevated associations with harmful or aggressive semantic domains. Many of these patterns align with traditional phonesthemes identified for English: /gl-/ → visual brightness (e.g., glimmer, glint, gleam), /sn-/ → nasal/oral actions (e.g., sniff, snort, sneeze), and /fl-/ → fluid motion (e.g., flow, flutter, flap) — all of them corroborated at scale with moderate effect sizes after Bonferroni correction.

Our findings validate the existence of distributed iconicity in everyday vocabulary, extending beyond classic effects into neurobiologically grounded experiential space. By juxtaposing the results with syntheses of behavioral phonosemantic research, our study confirms that computational profiling captures the perceptual-analogical mechanisms isolated in behavioral paradigms.

Contribution

The present work aims to showcase large-scale computational phonosemantics as a contribution to digital humanities research that integrates contemporary linguistic theory with advanced data analysis methods. By scaling experiential semantic annotation through BERT-based regression, it removes reliance on costly and limited human norming while maintaining strong generalization across wide vocabularies and across languages. Through high-dimensional phonological parsing and gradient boosting, the approach enables fine-grained data mining within complex linguistic feature spaces, allowing systematic effects to be detected beyond isolated phonesthemes. By quantifying phonosemantic patterns across neurobiologically grounded experiential dimensions and situating them within established psycholinguistic findings, our study links statistical regularities in lexical structure to embodied mechanisms of perception and cognition. In doing so, it advances an integrative view of language in which arbitrariness and iconicity jointly shape lexical organization and everyday meaning-making.

References
  1. Balota, David A. / Yap, Melvin J. / Hutchison, Keith A. / Cortese, Michael J. / Kessler, Brett / Loftis, Bjorn / Neely, James H. / Nelson, Douglas L. / Simpson, Greg B. / Treiman, Rebecca (2007): “The English Lexicon Project”, in: Behavior Research Methods 39: 445–459.
  2. Batsuren, Khuyagbaatar / Bella, Gábor / Giunchiglia, Fausto (2021): “MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology”, in: Proceedings of the 18th SIGMORPHON Workshop on Computational Research in Phonetics: 39–48.
  3. Binder, Jeffrey R. / Conant, Lisa L. / Humphries, Colin J. / Fernandino, Leonardo / Simons, Stephen B. / Aguilar, Mario / Desai, Rutvik H. (2016): “Toward a Brain-Based Componential Semantic Representation”, in: Cognitive Neuropsychology 33, 3–4: 130–174.
  4. Blasi, Damián E. / Wichmann, Søren / Hammarström, Harald / Stadler, Peter F. / Christiansen, Morten H. (2016): “Sound–meaning association biases evidenced across thousands of languages”, in: Proceedings of the National Academy of Sciences 113, 39: 10818–10823.
  5. Carnegie Mellon Pronouncing Dictionary [CMUDict] (2014): Carnegie Mellon University. <http://www.speech.cs.cmu.edu/> [07.05.2026].
  6. Chelba, Ciprian / Mikolov, Tomas / Schuster, Mike / Ge, Qi / Brants, Thorsten / Koehn, Phillipp / Robinson, Tony (2013): “One billion word benchmark for measuring progress in statistical language modeling.” <https://arxiv.org/abs/1312.3005> [07.05.2026].
  7. Devlin, Jacob / Chang, Ming-Wei / Lee, Kenton / Toutanova, Kristina (2019): “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, in: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT): 4171–4186.
  8. Dingemanse, Mark / Blasi, Damián E. / Lupyan, Gary / Christiansen, Morten H. / Monaghan, Padraic (2015): “Arbitrariness, iconicity, and systematicity in language”, in: Trends in Cognitive Sciences 19, 10: 603–615.
  9. Erben Johansson, Niklas / Anikin, Andrey / Carling, Gerd / Holmer, Arthur (2020): “The typology of sound symbolism: Defining macro-concepts via their semantic and phonetic features”, in: Linguistic Typology 24, 2: 253–310.
  10. Gutiérrez, Enrique D. / Levy, Roger / Bergen, Benjamin K. (2016): “Finding non-arbitrary form-meaning systematicity using stringmetric learning for kernel regression”, in: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers): 2379–2388.
  11. Lockwood, Gwilym / Dingemanse, Mark (2015): “Iconicity in the lab: A review of behavioral, developmental, and neuroimaging research into sound-symbolism”, in: Frontiers in Psychology 6: 1246.
  12. Meissel, Kane / Yao, Esther S. (2024): “Using Cliff’s delta as a non-parametric effect size measure: an accessible web app and R tutorial”, in: Practical Assessment, Research, and Evaluation 29, 1.
  13. Monaghan, Padraic / Shillcock, Richard C. / Christiansen, Morten H. / Kirby, Simon (2014): “How arbitrary is language?”, in: Philosophical Transactions of the Royal Society B: Biological Sciences 369, 1651.
  14. Mroczkowski, Robert / Rybak, Piotr / Wróblewska, Alina / Gawlik, Ireneusz (2021): “HerBERT: Efficiently Pretrained Transformer-based Language Model for Polish”, in: Proceedings of the 8th Workshop on Balto-Slavic Natural Language Processing: 1–10.
  15. Ngai, Chun Hau / Kilpatrick, Alexander J. / Ćwiek, Aleksandra (2024): “Sound symbolism in Japanese names: Machine learning approaches to gender classification”, in: PLOS One 19, 3: e0297440.
  16. Peirce, Charles Sanders (1931–1958): Collected Papers. Cambridge, MA: Harvard University Press.
  17. Pindur, Szymon (2025): “Sound symbolism in speakers of English: a qualitative synthesis”, in: International Journal of English Linguistics 15, 2.
  18. Plato (1961): “Cratylus”, in: Hamilton, Edith / Cairns, Huntington (eds.): The Collected Dialogues. Princeton, NJ: Princeton University Press: 421–476.
  19. Saussure, Ferdinand de (1916/1966): Course in General Linguistics. New York, NY: McGraw-Hill.
  20. Sidhu, David M. / Pexman, Penny M. (2018): “Five mechanisms of sound symbolic association”, in: Psychonomic Bulletin & Review 25: 1619–1643.
  21. Turton, Jacob / Vinson, David P. / Smith, Robert Elliott (2021): “Deriving contextualised semantic features from BERT (and other transformer model) embeddings”, in: Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021): 248–262.
  22. Winter, Bodo / Perlman, Marcus (2021): “Size sound symbolism in the English lexicon”, in: Glossa: a Journal of General Linguistics 6, 1: 79.
  23. Wojdyga, Grzegorz (2018): “Results of the PolEval 2018 Shared Task 3: Language Models”, in: Proceedings of the PolEval 2018 Workshop: 121–127.
  24. Ylönen, Tatu (2022): “Wiktextract: Wiktionary as Machine-Readable Structured Data”, in: Proceedings of the 13th Conference on Language Resources and Evaluation (LREC): 1317–1325.