DH 2026

Daejeon, July 27–31

Thu, July 3009:00–10:30S028103
Long Paper

Introducing Linguistic Asymmetry: Critical Approaches to Metadata, Multilingualism, and Equity in Digital Humanities Infrastructures

Elena Battaner Moro
Universidad Rey Juan Carlos, Spain · elena.battaner@urjc.es
Paul Spence
King's College London, UK · paul.spence@kcl.ac.uk

Digital infrastructures increasingly mediate access to cultural heritage, research data, and scholarly communication across world regions, and understanding how they function is a fundamental element of knowledge work (Bowker et al. 2009; Liu 2018; Pawlicka-Deger 2022). As these systems incorporate heterogeneous collections and multilingual materials, they also encode structural forms of linguistic imbalance (Nilsson-Fernàndez & Dombrowski 2022; Sneha & Sengupta 2021; Zaugg et al. 2022). Such imbalances emerge not from explicit exclusion but from design decisions within metadata schemas, aggregation pipelines, institutional hierarchies, and transliteration conventions. This paper proposes the concept of linguistic asymmetry as a global and systemic property of digital infrastructures and introduces the Linguistic Asymmetry Index (LAI) as a reproducible and FAIR-aligned diagnostic for analysing this condition across diverse settings.

Widespread references to “multilingual infrastructures” often equate multilingualism with the mere co-presence of multiple languages, yet multilingual presence does not guarantee linguistic equity. A platform may include resources in many languages while privileging a single metadata pivot — that is, the language to which descriptive fields default when records are aggregated; it may inconsistently populate language fields, apply uneven transliteration rules, or rely on controlled vocabularies tied to only one linguistic tradition. These practices affect visibility, retrieval, and interoperability in ways that disproportionately benefit some languages over others (Spence 2021; Viola 2024), and are particularly consequential where infrastructures mediate collections in dozens or hundreds of languages, display multiple writing systems (Horvath 2021), and operate under pressures to align metadata with English or another dominant lingua franca.

We therefore define linguistic asymmetry as an infrastructural condition characterised by unequal representation, completeness, discoverability, or access across languages. It emerges from technical normalisation, institutional concentration, or algorithmic mediation rather than deliberate linguistic policy, and reframes a familiar observation — the dominance of highly resourced languages in digital systems — as a problem of design, governance, and metadata epistemology. Following Bowker and Star’s account of classification as a site of structural inclusion and exclusion (1999), and in dialogue with Gitelman’s critique of “raw data” (2013) and Parry’s argument that metadata actively configure what digital objects are (2023), the LAI treats linguistic asymmetry as a structural consequence of documentation practices that shape the visibility and legibility of languages within digital systems. The argument is positioned within critical infrastructure studies, theories of metadata as epistemic mediation, and debates on equity in global digital humanities (Fiormonte, Chaudhuri and Ricaurte 2022).

To operationalise this concept, we have developed the LAI, a composite diagnostic across five components. (1) Language Representation Asymmetry (LRA) measures distributional imbalance using either the Gini coefficient

The Gini coefficient is widely used as a statistical measure of economic (in)equality.

— a 0–1 inequality index where 0 = perfect parity and 1 = total concentration — or the Herfindahl–Hirschman Index (HHI),

The Herfindahl–Hirschman Index (HHI) is commonly used to measure market concentration.

the sum of squared language shares. (2) English or Anchor-Language Bias (ALB) captures the dominance of a single metadata pivot relative to other languages. (3) Metadata Completeness Disparity (MCD) records variation in the population of key descriptive fields (typically title, description/abstract, subject/keywords, creator, publisher, language) across languages. (4) Institutional Concentration Index (ICI) measures the degree to which a small number of providers determine the linguistic landscape of the infrastructure. (5) Access Inequality Index (AII) records differences in licensing or openness affecting languages unequally. Each component is computed from public APIs or OAI-PMH endpoints; thresholds, vocabularies, and exclusions are documented per source.

Methodologically, the LAI aggregates these components through min–max normalisation followed by an equal-weighted arithmetic mean as the default specification, with all weights, thresholds, and metric choices explicitly declared. The framework is intentionally reweightable: institutions may downweight components irrelevant to their mandate (for example ‘AII’ for a closed-access archive that publishes its policy explicitly), substitute Gini for HHI in LRA, or extend MCD to additional fields. The full workflow, datasets, and preliminary results are openly available in Zenodo (Battaner & Spence 2025), and this approach has now been applied empirically across various complementary studies: a cross-infrastructure analysis covering CLARIN, EUDAT/B2FIND, Europeana, and OpenAire (Battaner & Spence 2026a), and federated audits of GoTriple and CLARIN’s VLO (Battaner et al. 2026; Battaner 2026 forthcoming, Candela et al. 2026).

Our pilot applications illustrate this variability and the explanatory value of separating components. CLARIN exhibits asymmetry linked to disciplinary conventions that favour English in key metadata fields, despite a heterogeneous languageCode field that nominally accommodates more than five thousand variant values. Europeana shows a pattern shaped by its distributed provider network, where multilingual coverage is extensive but organised around national collections, producing high LRA balance with strong ICI concentration. EUDAT’s uniform metadata model, which treats language as an optional field, results in limited linguistic differentiation: low ALB but also low MCD because the field is rarely populated at all. OpenAIRE reflects broader publication practices, with English predominance influencing its aggregated records. The federated audit of GoTriple (Battaner & al. 2026) further shows how aggregation itself reshapes the picture, with undefined and other language records absorbing minoritised production. Crucially, these patterns cannot be inferred from content statistics alone: they require analysing metadata as the locus where linguistic visibility is produced.

The LAI is conceived in dialogue with the Research Data Alliance’s work on metadata interoperability (2023) and the Ethical Data Initiative’s principles on data justice (2024), situating linguistic equity as a core dimension of open and reproducible research. By extending the FAIR principles toward language — what we and others have begun to call FAIR-by-language (Viola 2024) — the framework offers a vocabulary that connects technical metadata governance with broader open-science policy debates.

Two methodological caveats follow. First, the LAI is computed from data that infrastructures themselves expose; this raises a legitimate concern that an archive is, in effect, marking its own homework. Our response is structural: the LAI is run against the outward-facing surface of the infrastructure — public APIs, OAI-PMH endpoints, Solr facets — so what it measures is what the archive promises to its community, not what it claims internally. Triangulation with federated aggregators (GoTriple, OpenAIRE) and with parallel discovery layers makes deviations between declared and exposed metadata visible. Second, quantification cannot capture linguistic vitality, cultural significance, or historical context. The LAI is therefore not a ranking instrument but a diagnostic lens that identifies structural patterns requiring contextual reading.

This is especially important where infrastructures handle Indigenous, minority, or endangered languages. We therefore propose a three-step qualitative integration protocol: (i) prior consultation with linguistic custodians, archival experts, and community representatives before applying the LAI to such collections; (ii) suspension of quantitative interpretation wherever metadata coverage is too sparse or too inconsistent to support inference, in favour of contextual reading; and (iii) open publication of the contextual reading alongside the numerical score, not as an appendix but as a constitutive part of the output. Participatory adaptation of components is encouraged.

The contribution of the LAI lies in reframing infrastructural assessment as a critical diagnostic practice. Conventional benchmarking evaluates performance or compliance; the LAI is standardised in method but openly interpretive in reading. There is no contradiction between the two: a transparent, reproducible workflow is precisely what allows context-sensitive interpretation to be defended, contested, and refined across communities. The framework aligns with open, transparent and reproducible research principles and extends them toward linguistic visibility and infrastructural justice.

In sum, the paper contributes a conceptual vocabulary, a reproducible method, and an interpretive framework for analysing linguistic equity in digital infrastructures. By decentring specific regional cases and focusing on global infrastructural patterns, it invites the DH community to examine how its systems construct or obscure linguistic plurality, and offers a practical, adaptable instrument for supporting more equitable metadata governance across the digital humanities.

References
  1. Battaner Moro, E. (2026). languageCode Field Distribution in the CLARIN Virtual Language Observatory -- Complete Solr Facet Export, 25 March 2026 (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.19388972
  2. Battaner Moro, E., Candela, G., Rosiński, C., Spence, P., & Wołczuk, N. (2026). Multilingualism at the Semantic Layer: Metadata Evidence from the GoTriple Research Infrastructure. Zenodo. https://doi.org/10.5281/zenodo.19262149
  3. Battaner Moro, E., & Spence, P. (2025). Linguistic Asymmetry Index (LAI): Benchmarking multilingual research infrastructures (Version 1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.17597231
  4. Battaner Moro, E. and Spence, P. (2026) “The Linguistic Asymmetry Index (LAI): Benchmarking Equity in Multilingual Research Infrastructures”, in: Journal of Open Humanities Data, 12(1), p. 28. Available at: https://doi.org/10.5334/johd.474.
  5. Bowker, G.C., Baker, K., Millerand, F., Ribes, D. (2009). “Toward Information Infrastructure Studies: Ways of Knowing in a Networked Environment”, n: Hunsinger, J., Klastrup, L., Allen, M. (eds) International Handbook of Internet Research. Dordrecht: Springer. https://doi.org/10.1007/978-1-4020-9789-8_5
  6. Candela, G., Wołczuk, N., Rosiński, C., Battaner Moro, E., & Spence, P. (2026). FASCA Project - Pilot 2 dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.19731084
  7. Bowker, G. C., & Star, S. L. (1999). Sorting things out: Classification and its consequences. MIT Press.
  8. Ethical Data Initiative (2024). Principles for ethical data practice. https://ethicaldatainitiative.org/
  9. Fiormonte, D., Chaudhuri, S., & Ricaurte, P. (Eds.). (2022). Global debates in the digital humanities. University of Minnesota Press.
  10. Horvath, A. (2021). “Enhancing language inclusivity in digital humanities: Towards sensitivity and multilingualism”, in: Modern Languages Open, 1(26), 1–12. https://doi.org/10.3828/mlo.v0i0.382
  11. Liu, A. (2018). “Toward Critical Infrastructure Studies.” Critical Infrastructure Studies (CIstudies.Org) (blog). http://cistudies.org/wp-content/uploads/Toward-Critical-Infrastructure-Studies.pdf.
  12. Nilsson-Fernàndez, P. & Dombrowski, Q. (2022). “Multilingual Digital Humanities”, in: The Bloomsbury Handbook of Digital Humanities. London: Bloomsbury Academic..
  13. Parry, K. (2023). Metadata is not “data about data”, in: M. Filimowicz (Ed.), Decolonizing data: Algorithms and society (pp. 16–32). Routledge. https://doi.org/10.4324/9781003299912
  14. Pawlicka-Deger, U. (2022). “Infrastructuring digital humanities: On relational infrastructure and global reconfiguration of the field”, in: Digital Scholarship in the Humanities, Volume 37, Issue 2, pp. 534–550, https://doi.org/10.1093/llc/fqab086
  15. Research Data Alliance (2023). FAIR-aligned metadata interoperability: Recommendations and outputs. https://www.rd-alliance.org/
  16. Sneha, P. P., and Sengupta, A. (2021). "The Many Languages of Digital Infrastructures", in: India-seminar. https://www.india-seminar.com/2021/742/742_puthiya_and_anasuya.htm
  17. Spence, P. (2021). Disrupting digital monolingualism: A report on multilingualism in digital theory and practice (Version 1.0). Language Acts & Worldmaking Project. https://doi.org/10.5281/zenodo.5743283
  18. Viola, L. (2024). “Editorial: Data and workflows for multilingual digital humanities”, in: Journal of Open Humanities Data, 10(1), 37. https://doi.org/10.5334/johd.220
  19. Zaugg, I., Hossain, A. & Molloy, B. (2022). “Digitally-disadvantaged languages”, in: Internet Policy Review, 11(2). Available at: https://policyreview.info/glossary/digitally-disadvantaged-languages