DH 2026

Daejeon, July 27–31

Wed, July 2914:00–15:30S109108
Short Paper

Mapping Representational and Evaluative Meaning Across Socio-Semiotic Activities: A Digital Humanities Study of Potentially Idiomatic Expressions in Brazilian Portuguese

Adriana Pagano
Federal University of Minas Gerais, Brazil · apagano@ufmg.br
Ana Clara Pagano
Federal University of Minas Gerais, Brazil · anapagano.ufmg@gmail.com
Aline Villavicencio
University of Exeter · A.Villavicencio@exeter.ac.uk
Rodrigo Wilkens
University of Exeter · r.wilkens@exeter.ac.uk

Introduction

Potentially idiomatic expressions (PIEs) pose a persistent challenge for Natural Language Processing and Digital Humanities research. Their non-compositional and context-dependent meanings may lead to misclassification in computational pipelines, particularly in multilingual settings and in text types where stance-taking and evaluative density are high (Shutova 2011; Constant et al. 2017). At the same time, PIEs function as culturally embedded resources for interpersonal alignment, identity construction, and community-specific modes of expression (Savary et al. 2017). Recent benchmark-based work has advanced the cross-lingual and multimodal evaluation of idiomaticity (Torunoğlu-Selamet et al. 2026). However, less attention has been paid to how potentially idiomatic expressions vary across socio-semiotic activities and evaluative functions, particularly within less represented languages. While Digital Humanities scholarship increasingly engages with meaning-making at scale, existing work tends to focus either on metaphor more broadly (Underwood 2019; Piper 2020), on English-language corpora, or on single text type explorations. Research integrating text type variation, functional linguistics, phraseology, and corpus-based analysis remains limited.

This study addresses this gap by examining PIE usage in a little represented language  — Brazilian Portuguese  — across multiple socio-semiotic activities. In this study, PIEs are understood as multiword or recurrent expressions that have the potential to convey non-compositional meaning, but whose actual interpretation depends on context. For example, in Brazilian Portuguese the expression a cereja do bolo “the cherry on the cake” may be used idiomatically to mean the final or most valued element of a situation, but it may also occur literally in contexts referring to an actual cherry placed on a cake. Its classification therefore depends on co-text and domain.

The broader relevance of classifying PIEs extends beyond linguistic description. Since PIEs often encode stance, evaluation, cultural knowledge, and community-specific forms of expression, their automatic or semi-automatic identification can support large-scale studies of evaluative meaning, social positioning, and cultural variation in digital corpora. Improving the classification of PIEs in Brazilian Portuguese is therefore important not only for idiom studies, but also for multilingual NLP, corpus-based Digital Humanities, sentiment analysis, stance detection, and the study of cultural meaning-making across text types.

Theoretical Framework

The study combines Systemic Functional Linguistics, Appraisal Theory, and corpus-based semantic analysis. From a Systemic Functional Linguistics perspective, PIEs are treated as resources that contribute both to representational meaning — how experience is construed — and to interpersonal meaning — how speakers and writers evaluate people, events, situations, and propositions (Halliday / Matthiessen 2014). Within Appraisal Theory, the analysis focuses on Attitude and Graduation (Martin / White 2005). Attitude is examined through three evaluative dimensions: Affect, which concerns emotions; Judgement, which concerns evaluations of people and behaviour; and Appreciation, which concerns evaluations of things, events, and states of affairs. Graduation captures differences in evaluative force, intensity, and categorical sharpening or softening.

Although Appraisal Theory also includes Engagement, the present analysis focuses on Attitude and Graduation. Engagement was excluded from the final quantitative annotation because dialogic positioning often depends on broader discourse context rather than on the PIE occurrence itself, especially in literary dialogue, reported speech, and social media interaction.

To situate PIEs in their communicative environments, the study adopts Halliday and Matthiessen’s typology of socio-semiotic activities (Halliday / Matthiessen 2014). The corpus represents four activities: reporting, expounding, sharing/exploring, and recreating. These activities provide a functional basis for comparing how PIEs vary across informational, interpersonal, and narrative contexts.

Corpus and Methodology

The corpus comprises Brazilian Portuguese texts representing four socio-semiotic activities: news reports, popular science articles, Reddit social media posts, and contemporary literary narratives. These domains were selected to allow comparison across edited informational discourse, everyday digital interaction, and literary meaning-making.

The analysis began with a master list of PIE candidates compiled from Brazilian Portuguese idiomatic and phraseological expressions. This list was used as a lexicon for corpus extraction through a Python script. Candidate occurrences were identified through string-based matching, followed by manual checking of concordance lines to verify whether each occurrence corresponded to a relevant PIE use. Each occurrence was then manually classified according to its representational status as either idiomatic or literal. In addition, each occurrence was annotated for Appraisal categories: Affect, Judgement, Appreciation, and Graduation. This procedure made it possible to compare both the distribution of PIEs across domains and their evaluative functions in context.

Results

The analysis identified 442 PIE occurrences across the corpus, corresponding to 112 distinct PIE types. Idiomatic uses strongly predominate overall: 328 occurrences were classified as idiomatic and 114 as literal, meaning that approximately 74% of all PIE tokens are idiomatic. However, the distribution varies considerably across domains. Social media contains the highest number of PIE occurrences (167) and is overwhelmingly idiomatic, with 166 idiomatic uses and only 1 literal use. Science popularization follows with 127 occurrences, of which 92 are idiomatic and 35 literal. News reports contain 81 occurrences, with 51 idiomatic and 30 literal uses. Literary texts show a different pattern: although they contain 67 occurrences, only 19 are idiomatic, while 48 are literal.

Type diversity follows a similar domain-sensitive distribution. Social media contains the largest number of distinct PIEs (64 types), followed by news (39), science popularization (36), and literary texts (16). These findings suggest that PIE behaviour is strongly shaped by socio-semiotic activity: social media favours idiomatic and phraseologically diverse uses associated with interpersonal stance-taking, while edited informational and literary domains preserve a higher proportion of literal or compositional readings.

The Appraisal annotation further shows that PIEs are strongly associated with evaluative meaning, although the relevant evaluative dimensions vary across domains. Affect is absent in many occurrences overall, but when present it is frequently associated with negative emotions such as dissatisfaction and unhappiness. Judgement appears unevenly across domains, with social media showing a strong concentration of negative evaluations of conduct, especially negative social sanction and propriety. Appreciation is the most consistently activated Attitude category, with negative valuation accounting for a large share of the data overall. This suggests that PIEs are frequently used to evaluate situations, events, and states of affairs as problematic, undesirable, or noteworthy. Graduation patterns reinforce this interpretation: focus sharpening and force intensification are frequent resources, showing that PIEs often function not only to evaluate experience but also to intensify stance and sharpen category boundaries.

Contribution

This study makes three main contributions. First, it offers a theoretically grounded and corpus-assisted method for studying PIEs across domains, combining Systemic Functional Linguistics, Appraisal Theory, and corpus-based extraction. Second, it expands Digital Humanities and NLP research on idiomaticity by focusing on Brazilian Portuguese, a major world language that remains underrepresented in computational and DH studies of evaluative meaning. Third, it provides a replicable workflow for identifying, classifying, and functionally annotating PIEs in multi-genre corpora.

Overall, the project demonstrates that combining computational methods with functional linguistic theory enables a more nuanced analysis of idiomaticity, stance, and cultural expression. By examining PIEs across socio-semiotic activities, the study shows how representational and evaluative meanings circulate through contemporary cultural domains, from news and popular science to social media and literary narrative.

References
  1. Constant, Mathieu / Sigogne, Anthony / Watrin, Patrick (2017): “Multiword expression processing: A survey of the last decade and future directions”, in: Computational Linguistics 43, 4: 837–892. DOI: 10.1162/COLI_a_00302.
  2. Halliday, M. A. K. / Matthiessen, Christian M. I. M. (2014): Halliday’s Introduction to Functional Grammar. 4th ed. London / New York: Routledge.
  3. Martin, J. R. / White, P. R. R. (2005): The Language of Evaluation: Appraisal in English. London: Palgrave Macmillan.
  4. Piper, Andrew (2020): Enumerations: Data and Literary Study. Chicago: University of Chicago Press.
  5. Sag, Ivan A. et al. (2002): “Multiword expressions: A pain in the neck for NLP”, in: Gelbukh, Alexander (ed.): Computational Linguistics and Intelligent Text Processing. CICLing 2002. Berlin / Heidelberg: Springer 1–15.
  6. Savary, Agata et al. (2017): “The PARSEME shared task on automatic identification of verbal multiword expressions”, in: Proceedings of the 13th Workshop on Multiword Expressions. Valencia: Association for Computational Linguistics 31–47.
  7. Shutova, Ekaterina (2011): “Design and evaluation of metaphor processing systems”, in: Computational Linguistics 39, 1: 1–28. DOI: 10.1162/coli_a_00033.
  8. Tagnin, Stella E. O. (2013): O jeito que a gente diz: expressões convencionais e idiomáticas. São Paulo: Disal.
  9. Torunoğlu-Selamet, Dilara et al. (2026): “A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding”, in: Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). Palma, Mallorca, Spain: European Language Resources Association (ELRA) 9434–9448. DOI: 10.63317/5cvnbcoktfo2.
  10. Underwood, Ted (2019): Distant Horizons: Digital Evidence and Literary Change. Chicago: University of Chicago Press.
  11. Zappavigna, Michele (2012): Discourse of Twitter and Social Media: How We Use Language to Create Affiliation on the Web. London / New York: Bloomsbury Academic.