Daejeon, July 27–31
Potentially idiomatic expressions (PIEs) pose a persistent challenge for Natural Language Processing and Digital Humanities research. Their non-compositional and context-dependent meanings may lead to misclassification in computational pipelines, particularly in multilingual settings and in text types where stance-taking and evaluative density are high (Shutova 2011; Constant et al. 2017). At the same time, PIEs function as culturally embedded resources for interpersonal alignment, identity construction, and community-specific modes of expression (Savary et al. 2017). Recent benchmark-based work has advanced the cross-lingual and multimodal evaluation of idiomaticity (Torunoğlu-Selamet et al. 2026). However, less attention has been paid to how potentially idiomatic expressions vary across socio-semiotic activities and evaluative functions, particularly within less represented languages. While Digital Humanities scholarship increasingly engages with meaning-making at scale, existing work tends to focus either on metaphor more broadly (Underwood 2019; Piper 2020), on English-language corpora, or on single text type explorations. Research integrating text type variation, functional linguistics, phraseology, and corpus-based analysis remains limited.
This study addresses this gap by examining PIE usage in a little represented language — Brazilian Portuguese — across multiple socio-semiotic activities. In this study, PIEs are understood as multiword or recurrent expressions that have the potential to convey non-compositional meaning, but whose actual interpretation depends on context. For example, in Brazilian Portuguese the expression a cereja do bolo “the cherry on the cake” may be used idiomatically to mean the final or most valued element of a situation, but it may also occur literally in contexts referring to an actual cherry placed on a cake. Its classification therefore depends on co-text and domain.
The broader relevance of classifying PIEs extends beyond linguistic description. Since PIEs often encode stance, evaluation, cultural knowledge, and community-specific forms of expression, their automatic or semi-automatic identification can support large-scale studies of evaluative meaning, social positioning, and cultural variation in digital corpora. Improving the classification of PIEs in Brazilian Portuguese is therefore important not only for idiom studies, but also for multilingual NLP, corpus-based Digital Humanities, sentiment analysis, stance detection, and the study of cultural meaning-making across text types.
The study combines Systemic Functional Linguistics, Appraisal Theory, and corpus-based semantic analysis. From a Systemic Functional Linguistics perspective, PIEs are treated as resources that contribute both to representational meaning — how experience is construed — and to interpersonal meaning — how speakers and writers evaluate people, events, situations, and propositions (Halliday / Matthiessen 2014). Within Appraisal Theory, the analysis focuses on Attitude and Graduation (Martin / White 2005). Attitude is examined through three evaluative dimensions: Affect, which concerns emotions; Judgement, which concerns evaluations of people and behaviour; and Appreciation, which concerns evaluations of things, events, and states of affairs. Graduation captures differences in evaluative force, intensity, and categorical sharpening or softening.
Although Appraisal Theory also includes Engagement, the present analysis focuses on Attitude and Graduation. Engagement was excluded from the final quantitative annotation because dialogic positioning often depends on broader discourse context rather than on the PIE occurrence itself, especially in literary dialogue, reported speech, and social media interaction.
To situate PIEs in their communicative environments, the study adopts Halliday and Matthiessen’s typology of socio-semiotic activities (Halliday / Matthiessen 2014). The corpus represents four activities: reporting, expounding, sharing/exploring, and recreating. These activities provide a functional basis for comparing how PIEs vary across informational, interpersonal, and narrative contexts.
The corpus comprises Brazilian Portuguese texts representing four socio-semiotic activities: news reports, popular science articles, Reddit social media posts, and contemporary literary narratives. These domains were selected to allow comparison across edited informational discourse, everyday digital interaction, and literary meaning-making.
The analysis began with a master list of PIE candidates compiled from Brazilian Portuguese idiomatic and phraseological expressions. This list was used as a lexicon for corpus extraction through a Python script. Candidate occurrences were identified through string-based matching, followed by manual checking of concordance lines to verify whether each occurrence corresponded to a relevant PIE use. Each occurrence was then manually classified according to its representational status as either idiomatic or literal. In addition, each occurrence was annotated for Appraisal categories: Affect, Judgement, Appreciation, and Graduation. This procedure made it possible to compare both the distribution of PIEs across domains and their evaluative functions in context.
The analysis identified 442 PIE occurrences across the corpus, corresponding to 112 distinct PIE types. Idiomatic uses strongly predominate overall: 328 occurrences were classified as idiomatic and 114 as literal, meaning that approximately 74% of all PIE tokens are idiomatic. However, the distribution varies considerably across domains. Social media contains the highest number of PIE occurrences (167) and is overwhelmingly idiomatic, with 166 idiomatic uses and only 1 literal use. Science popularization follows with 127 occurrences, of which 92 are idiomatic and 35 literal. News reports contain 81 occurrences, with 51 idiomatic and 30 literal uses. Literary texts show a different pattern: although they contain 67 occurrences, only 19 are idiomatic, while 48 are literal.
Type diversity follows a similar domain-sensitive distribution. Social media contains the largest number of distinct PIEs (64 types), followed by news (39), science popularization (36), and literary texts (16). These findings suggest that PIE behaviour is strongly shaped by socio-semiotic activity: social media favours idiomatic and phraseologically diverse uses associated with interpersonal stance-taking, while edited informational and literary domains preserve a higher proportion of literal or compositional readings.
The Appraisal annotation further shows that PIEs are strongly associated with evaluative meaning, although the relevant evaluative dimensions vary across domains. Affect is absent in many occurrences overall, but when present it is frequently associated with negative emotions such as dissatisfaction and unhappiness. Judgement appears unevenly across domains, with social media showing a strong concentration of negative evaluations of conduct, especially negative social sanction and propriety. Appreciation is the most consistently activated Attitude category, with negative valuation accounting for a large share of the data overall. This suggests that PIEs are frequently used to evaluate situations, events, and states of affairs as problematic, undesirable, or noteworthy. Graduation patterns reinforce this interpretation: focus sharpening and force intensification are frequent resources, showing that PIEs often function not only to evaluate experience but also to intensify stance and sharpen category boundaries.
This study makes three main contributions. First, it offers a theoretically grounded and corpus-assisted method for studying PIEs across domains, combining Systemic Functional Linguistics, Appraisal Theory, and corpus-based extraction. Second, it expands Digital Humanities and NLP research on idiomaticity by focusing on Brazilian Portuguese, a major world language that remains underrepresented in computational and DH studies of evaluative meaning. Third, it provides a replicable workflow for identifying, classifying, and functionally annotating PIEs in multi-genre corpora.
Overall, the project demonstrates that combining computational methods with functional linguistic theory enables a more nuanced analysis of idiomaticity, stance, and cultural expression. By examining PIEs across socio-semiotic activities, the study shows how representational and evaluative meanings circulate through contemporary cultural domains, from news and popular science to social media and literary narrative.