Daejeon, July 27–31
Digital infrastructures increasingly mediate access to cultural heritage, research data, and scholarly communication across world regions, and understanding how they function is a fundamental element of knowledge work (Bowker et al. 2009; Liu 2018; Pawlicka-Deger 2022). As these systems incorporate heterogeneous collections and multilingual materials, they also encode structural forms of linguistic imbalance (Nilsson-Fernàndez & Dombrowski 2022; Sneha & Sengupta 2021; Zaugg et al. 2022). Such imbalances emerge not from explicit exclusion but from design decisions within metadata schemas, aggregation pipelines, institutional hierarchies, and transliteration conventions. This paper proposes the concept of linguistic asymmetry as a global and systemic property of digital infrastructures and introduces the Linguistic Asymmetry Index (LAI) as a reproducible and FAIR-aligned diagnostic for analysing this condition across diverse settings.
Widespread references to “multilingual infrastructures” often equate multilingualism with the mere co-presence of multiple languages, yet multilingual presence does not guarantee linguistic equity. A platform may include resources in many languages while privileging a single metadata pivot — that is, the language to which descriptive fields default when records are aggregated; it may inconsistently populate language fields, apply uneven transliteration rules, or rely on controlled vocabularies tied to only one linguistic tradition. These practices affect visibility, retrieval, and interoperability in ways that disproportionately benefit some languages over others (Spence 2021; Viola 2024), and are particularly consequential where infrastructures mediate collections in dozens or hundreds of languages, display multiple writing systems (Horvath 2021), and operate under pressures to align metadata with English or another dominant lingua franca.
We therefore define linguistic asymmetry as an infrastructural condition characterised by unequal representation, completeness, discoverability, or access across languages. It emerges from technical normalisation, institutional concentration, or algorithmic mediation rather than deliberate linguistic policy, and reframes a familiar observation — the dominance of highly resourced languages in digital systems — as a problem of design, governance, and metadata epistemology. Following Bowker and Star’s account of classification as a site of structural inclusion and exclusion (1999), and in dialogue with Gitelman’s critique of “raw data” (2013) and Parry’s argument that metadata actively configure what digital objects are (2023), the LAI treats linguistic asymmetry as a structural consequence of documentation practices that shape the visibility and legibility of languages within digital systems. The argument is positioned within critical infrastructure studies, theories of metadata as epistemic mediation, and debates on equity in global digital humanities (Fiormonte, Chaudhuri and Ricaurte 2022).
To operationalise this concept, we have developed the LAI, a composite diagnostic across five components. (1) Language Representation Asymmetry (LRA) measures distributional imbalance using either the Gini coefficient The Gini coefficient is widely used as a statistical measure of economic (in)equality. The Herfindahl–Hirschman Index (HHI) is commonly used to measure market concentration.
Methodologically, the LAI aggregates these components through min–max normalisation followed by an equal-weighted arithmetic mean as the default specification, with all weights, thresholds, and metric choices explicitly declared. The framework is intentionally reweightable: institutions may downweight components irrelevant to their mandate (for example ‘AII’ for a closed-access archive that publishes its policy explicitly), substitute Gini for HHI in LRA, or extend MCD to additional fields. The full workflow, datasets, and preliminary results are openly available in Zenodo (Battaner & Spence 2025), and this approach has now been applied empirically across various complementary studies: a cross-infrastructure analysis covering CLARIN, EUDAT/B2FIND, Europeana, and OpenAire (Battaner & Spence 2026a), and federated audits of GoTriple and CLARIN’s VLO (Battaner et al. 2026; Battaner 2026 forthcoming, Candela et al. 2026).
Our pilot applications illustrate this variability and the explanatory value of separating components. CLARIN exhibits asymmetry linked to disciplinary conventions that favour English in key metadata fields, despite a heterogeneous languageCode field that nominally accommodates more than five thousand variant values. Europeana shows a pattern shaped by its distributed provider network, where multilingual coverage is extensive but organised around national collections, producing high LRA balance with strong ICI concentration. EUDAT’s uniform metadata model, which treats language as an optional field, results in limited linguistic differentiation: low ALB but also low MCD because the field is rarely populated at all. OpenAIRE reflects broader publication practices, with English predominance influencing its aggregated records. The federated audit of GoTriple (Battaner & al. 2026) further shows how aggregation itself reshapes the picture, with undefined and other language records absorbing minoritised production. Crucially, these patterns cannot be inferred from content statistics alone: they require analysing metadata as the locus where linguistic visibility is produced.
The LAI is conceived in dialogue with the Research Data Alliance’s work on metadata interoperability (2023) and the Ethical Data Initiative’s principles on data justice (2024), situating linguistic equity as a core dimension of open and reproducible research. By extending the FAIR principles toward language — what we and others have begun to call FAIR-by-language (Viola 2024) — the framework offers a vocabulary that connects technical metadata governance with broader open-science policy debates.
Two methodological caveats follow. First, the LAI is computed from data that infrastructures themselves expose; this raises a legitimate concern that an archive is, in effect, marking its own homework. Our response is structural: the LAI is run against the outward-facing surface of the infrastructure — public APIs, OAI-PMH endpoints, Solr facets — so what it measures is what the archive promises to its community, not what it claims internally. Triangulation with federated aggregators (GoTriple, OpenAIRE) and with parallel discovery layers makes deviations between declared and exposed metadata visible. Second, quantification cannot capture linguistic vitality, cultural significance, or historical context. The LAI is therefore not a ranking instrument but a diagnostic lens that identifies structural patterns requiring contextual reading.
This is especially important where infrastructures handle Indigenous, minority, or endangered languages. We therefore propose a three-step qualitative integration protocol: (i) prior consultation with linguistic custodians, archival experts, and community representatives before applying the LAI to such collections; (ii) suspension of quantitative interpretation wherever metadata coverage is too sparse or too inconsistent to support inference, in favour of contextual reading; and (iii) open publication of the contextual reading alongside the numerical score, not as an appendix but as a constitutive part of the output. Participatory adaptation of components is encouraged.
The contribution of the LAI lies in reframing infrastructural assessment as a critical diagnostic practice. Conventional benchmarking evaluates performance or compliance; the LAI is standardised in method but openly interpretive in reading. There is no contradiction between the two: a transparent, reproducible workflow is precisely what allows context-sensitive interpretation to be defended, contested, and refined across communities. The framework aligns with open, transparent and reproducible research principles and extends them toward linguistic visibility and infrastructural justice.
In sum, the paper contributes a conceptual vocabulary, a reproducible method, and an interpretive framework for analysing linguistic equity in digital infrastructures. By decentring specific regional cases and focusing on global infrastructural patterns, it invites the DH community to examine how its systems construct or obscure linguistic plurality, and offers a practical, adaptable instrument for supporting more equitable metadata governance across the digital humanities.