Daejeon, July 27–31
“...There is, of course, the danger that the term sociotechnical system very rapidly becomes a shibboleth, the mere pronouncing of which distinguishes the cognoscenti from the ignorant and uninitiated.”
- Albert Cherns (The SocioTechnical Systems, 1971)
The digital divide in language technologies such as Machine Translation (MT) or Large Language Models (LLMs) reflects a systemic global inequity. This geographic and linguistic disparity is not merely a technical problem of data scarcity but stems from fundamental structural and ethical asymmetries in how NLP systems are designed and deployed across the majority worlds. Datasets, models, and benchmarks are doled out to language communities roughly in proportion to the revenue which their speakers can generate, while leaving the residual, under-served, or ‘low-resource’ categories disenfranchised (Nigatu et. al., 2024; Faisal et al., 2022). Furthermore, AI projects have often amounted to a rhetorical borrowing of social-science vocabulary such as participation, co-design, and community, with stakeholder involvement remaining tokenistic or extractive rather than substantively redistributive or at the least, exploratory (Sloane et al., 2020; Birhane et al., 2022; Delgado et al., 2023). Despite these faultlines, however, notable works have exhibited that a bottom-up, relational or participatory approach to MT/AI development, particularly for low-resource and endangered languages, can be adopted to address these challenges while centering the agency of native speakers (Nekoto et al., 2020; Bird, 2024). This paper introduces RAIL (Responsible Augmentation of Indigenous Languages), a framework for participatory data collection and annotation developed collaboratively with Marwari-speaking communities in Jodhpur villages. Marwari, a western Indo-Aryan language belonging to the Rajasthani language family, has about 7.8 lakh reported speakers spread across Jodhpur, Pali, Jaisalmer, Barmer, Nagaur, and Bikaner (Census of India, 2011). This approach aims to collate the linguistic expertise and cultural knowledge from native speakers to co-design the technology that they need.
Building on established approaches in community-based participatory research and principles articulated in research on participatory AI development (Šuster et al., 2023; Ovadya et al., 2024), RAIL adapts a five-tier engagement model: inform, consult, involve, collaborate, and empower to ensure that community members are active co-creators and decision-makers in the design, curation and governance of language technology assets (IAP2, 2018; Vaughn & Jacquez, 2020). It prefers meaningful process, attending to how technology is developed while abjuring any mechanistic preferences that may prioritize developing technological artifacts, because process matters more than artifacts in determining project success and community benefit (Cooper et al., 2024). A unique component of this framework comes through lessons from bioregionalism that deconstructs and debases the techno-positivist ontologies, which, by design, assume native speakers as mere data sources/points. Bioregionalism is a philosophy and practice that grounds cultural, economic, and political life in the ecological characteristics of a specific region (such as its watersheds and biomes), attending to the relationships between a place's environmental systems and the human communities and traditions situated within it (Berg & Dassmann, 1978). By attending to how a language is bound to the land, livelihoods, traditions, and ritual practices of the communities that carry it, this bioregional configuration to MT/AI development seeks to understand and preserve, rather than extract from, the linguistic worlds it engages.
To focus on governance, where most participatory projects tend to quietly revert to extraction, the framework learns from the CARE principles of Indigenous data sovereignty and community authority (Research Data Alliance, 2021). A notarized declaration is proposed to minimize legal friction and develop long-term relationships between the community and technologists. It proposes a path towards a local language body before sharing any derived artifacts (MT models, datasets, benchmarks) or enabling commercial uses. Selected members of the local language body, chosen based on social presence and community recognition, linguistic awareness, and demonstrated commitment to cultural preservation, retain veto power over data uses and derivative technologies. This approach is crucial to ensure sustainable community interests. This structure also operationalizes commitments to community collective benefit, responsibility for capacity building, and ethical protocols that respect community values (Research Data Alliance, 2021).
Recognizing the sustainability challenges inherent in participatory work, this framework, as a preliminary, also envisages a native-friendly technical repertoire of handbooks, articles and jupyter notebooks in local languages to facilitate a verifiable governance of digital assets and infrastructure supporting such projects. It also minimises recurring costs on volunteer training and data security dependencies on cloud providers. Findings also highlight that resources created during and from this process, enhance the capacity of native speakers to engage in larger global digital markets and discourses, creating sustainable economic incentives and livelihoods while motivating continued engagement with language technology development. This approach, besides other issues, also aims to address the major problem of geographical misrepresentation in NLP (Faisal et al., 2022). This work, besides the framework, showcases ethnographic insights on how relational NLP systems can be reproduced to serve community interests while respecting linguistic autonomy and data sovereignty in dissimilar geographic and social contexts. Rather than treating participatory involvement as a replacement to conventional NLP development, this interdisciplinary framework is positioned as a community-centered augmentative ad-on to enable successful community engagement and robust AI systems. Despite limitations in its current capacity to offer a wider representativeness or a shorter path, this framework seeks to pave the way for inclusive AI missions in India and beyond.
Based on questionnaires co-designed with inputs from native speakers, language experts and bioregionalism practitioners, 53 Marwari-speakers were engaged across 9 geographic locations in Jodhpur to develop an inter-generational learning repository which could also serve as a benchmark for AI evaluations. Till date, this engagement has produced approximately 400-800 annotated sentences per location with evidences of stark linguistic variances due to inter-community exclusivity. In addition to this, workshop material, visual booklets, and jupyter notebooks were developed to enhance communication with the participants. Hence, instead of aggregating language data in decontextualized datasets, RAIL grounds Marwari linguistic data within specific geographically, culturally, and socially relevant locations across Jodhpur, enabling the preservation of regional linguistic variation, dialectal features, and locally-situated language use patterns that are typically erased in massively multilingual datasets developed from geographically undifferentiated corpora (Bird et al. 2024, Faisal et al., 2022; Hovy & Purschke, 2019). For instance, community consultations have highlighted how language varieties associated with specific communities, such as linguistic patterns distinctive to the ‘Bishnoi’ community, differ from language use in other communities, living adjacently, such as the ‘Od’ or ‘Aanjna Patels’, for over hundreds of years. Such findings, aver that bioregional and community-specific data collection has potential to reveal linguistic nuances which top-down geography-blind datasets cannot capture.