DH 2026

Daejeon, July 27–31

Fri, July 3114:00–15:30S051103
Short Paper

Introducing RAIL: A Participatory Framework for Responsible AI

Himanshu Sihag
IIT Jodhpur, India · p23idh001@iitj.ac.in
“...There is, of course, the danger that the term sociotechnical system very rapidly becomes a shibboleth, the mere pronouncing of which distinguishes the cognoscenti from the ignorant and uninitiated.”
- Albert Cherns (The SocioTechnical Systems, 1971)

The digital divide in language technologies such as Machine Translation (MT) or Large Language Models (LLMs) reflects a systemic global inequity. This geographic and linguistic disparity is not merely a technical problem of data scarcity but stems from fundamental structural and ethical asymmetries in how NLP systems are designed and deployed across the majority worlds. Datasets, models, and benchmarks are doled out to language communities roughly in proportion to the revenue which their speakers can generate, while leaving the residual, under-served, or ‘low-resource’ categories disenfranchised (Nigatu et. al., 2024; Faisal et al., 2022). Furthermore, AI projects have often amounted to a rhetorical borrowing of social-science vocabulary such as participationco-design, and community, with stakeholder involvement remaining tokenistic or extractive rather than substantively redistributive or at the least, exploratory (Sloane et al., 2020; Birhane et al., 2022; Delgado et al., 2023). Despite these faultlines, however, notable works have exhibited that a bottom-up, relational or participatory approach to MT/AI development, particularly for low-resource and endangered languages, can be adopted to address these challenges while centering the agency of native speakers (Nekoto et al., 2020; Bird, 2024). This paper introduces RAIL (Responsible Augmentation of Indigenous Languages), a framework for participatory data collection and annotation developed collaboratively with Marwari-speaking communities in Jodhpur villages. Marwari, a western Indo-Aryan language belonging to the Rajasthani language family, has about 7.8 lakh reported speakers spread across Jodhpur, Pali, Jaisalmer, Barmer, Nagaur, and Bikaner (Census of India, 2011). This approach aims to collate the linguistic expertise and cultural knowledge from native speakers to co-design the technology that they need.

Building on established approaches in community-based participatory research and principles articulated in research on participatory AI development (Šuster et al., 2023; Ovadya et al., 2024), RAIL adapts a five-tier engagement model: inform, consult, involve, collaborate, and empower to ensure that community members are active co-creators and decision-makers in the design, curation and governance of language technology assets (IAP2, 2018; Vaughn & Jacquez, 2020). It prefers meaningful process, attending to how technology is developed while abjuring any mechanistic preferences that may prioritize developing technological artifacts, because process matters more than artifacts in determining project success and community benefit (Cooper et al., 2024). A unique component of this framework comes through lessons from bioregionalism that deconstructs and debases the techno-positivist ontologies, which, by design, assume native speakers as mere data sources/points. Bioregionalism is a philosophy and practice that grounds cultural, economic, and political life in the ecological characteristics of a specific region (such as its watersheds and biomes), attending to the relationships between a place's environmental systems and the human communities and traditions situated within it (Berg & Dassmann, 1978). By attending to how a language is bound to the land, livelihoods, traditions, and ritual practices of the communities that carry it, this bioregional configuration to MT/AI development seeks to understand and preserve, rather than extract from, the linguistic worlds it engages.

To focus on governance, where most participatory projects tend to quietly revert to extraction, the framework learns from the CARE principles of Indigenous data sovereignty and community authority (Research Data Alliance, 2021). A notarized declaration is proposed to minimize legal friction and develop long-term relationships between the community and technologists. It proposes a path towards a local language body before sharing any derived artifacts (MT models, datasets, benchmarks) or enabling commercial uses. Selected members of the local language body, chosen based on social presence and community recognition, linguistic awareness, and demonstrated commitment to cultural preservation, retain veto power over data uses and derivative technologies. This approach is crucial to ensure sustainable community interests. This structure also operationalizes commitments to community collective benefit, responsibility for capacity building, and ethical protocols that respect community values (Research Data Alliance, 2021).

Recognizing the sustainability challenges inherent in participatory work, this framework, as a preliminary, also envisages a native-friendly technical repertoire of handbooks, articles and jupyter notebooks in local languages to facilitate a verifiable governance of digital assets and infrastructure supporting such projects. It also minimises recurring costs on volunteer training and data security dependencies on cloud providers. Findings also highlight that resources created during and from this process, enhance the capacity of native speakers to engage in larger global digital markets and discourses, creating sustainable economic incentives and livelihoods while motivating continued engagement with language technology development. This approach, besides other issues, also aims to address the major problem of geographical misrepresentation in NLP (Faisal et al., 2022). This work, besides the framework, showcases ethnographic insights on how relational NLP systems can be reproduced to serve community interests while respecting linguistic autonomy and data sovereignty in dissimilar geographic and social contexts. Rather than treating participatory involvement as a replacement to conventional NLP development, this interdisciplinary framework is positioned as a community-centered augmentative ad-on to enable successful community engagement and robust AI systems. Despite limitations in its current capacity to offer a wider representativeness or a shorter path, this framework seeks to pave the way for inclusive AI missions in India and beyond.

Based on questionnaires co-designed with inputs from native speakers, language experts and bioregionalism practitioners, 53 Marwari-speakers were engaged across 9 geographic locations in Jodhpur to develop an inter-generational learning repository which could also serve as a benchmark for AI evaluations. Till date, this engagement has produced approximately 400-800 annotated sentences per location with evidences of stark linguistic variances due to inter-community exclusivity. In addition to this, workshop material, visual booklets, and jupyter notebooks were developed to enhance communication with the participants. Hence, instead of aggregating language data in decontextualized datasets, RAIL grounds Marwari linguistic data within specific geographically, culturally, and socially relevant locations across Jodhpur, enabling the preservation of regional linguistic variation, dialectal features, and locally-situated language use patterns that are typically erased in massively multilingual datasets developed from geographically undifferentiated corpora (Bird et al. 2024, Faisal et al., 2022; Hovy & Purschke, 2019). For instance, community consultations have highlighted how language varieties associated with specific communities, such as linguistic patterns distinctive to the ‘Bishnoi’ community, differ from language use in other communities, living adjacently, such as the ‘Od’ or ‘Aanjna Patels’, for over hundreds of years. Such findings, aver that bioregional and community-specific data collection has potential to reveal linguistic nuances which top-down geography-blind datasets cannot capture.

References
  1. Berg, P., & Dasmann, R. (1978). Reinhabiting a separate country: A bioregional anthology of Northern California (pp. 217–220). Planet Drum Foundation.
  2. Bird, S. (2024). Must NLP be extractive? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 14915–14929). Association for Computational Linguistics.
  3. Birhane, A., Isaac, W., Prabhakaran, V., Díaz, M., Elish, M. C., Gabriel, I., & Mohamed, S. (2022). Power to the people? Opportunities and challenges for participatory AI. In Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO ’22) (Article 6, pp. 1–8). Association for Computing Machinery. https://doi.org/10.1145/3551624.3555290
  4. Cherns, A. (1976). The principles of sociotechnical design. Human Relations, 29(8), 783–792.
  5. Cooper, N., Heldreth, C., & Hutchinson, B. (2024). “It’s how you do things that matters”: Attending to process to better serve Indigenous communities with language technologies. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 181–188).
  6. Delgado, F., Yang, S., Madaio, M., & Yang, Q. (2023). The participatory turn in AI design: Theoretical foundations and the current state of practice. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO ’23) (Article 37, pp. 1–23). Association for Computing Machinery. https://doi.org/10.1145/3617694.3623261
  7. Faisal, F., Wang, Y., & Anastasopoulos, A. (2022). Dataset geography: Mapping language data to language users. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics (pp. 1894–1907).
  8. Hovy, D., & Purschke, C. (2019). Exploring the linguistic landscape of geotagged social media data. Digital Scholarship in the Humanities, 34(2), 290–313.
  9. International Association for Public Participation. (2018). IAP2 spectrum of public participation. IAP2 Federation. https://www.iap2.org/page/pillars
  10. Madhavan, A., et al. (2025). Natural language processing applications for low-resource languages. Natural Language Processing, 1(1), 1–25.
  11. Nekoto, W., et al. (2020). Participatory research for low-resourced machine translation: A case study in African languages. In Findings of the Association for Computational Linguistics: EMNLP 2020 (pp. 2144–2160).
  12. Nigatu, H. H., Tonja, A. L., Rosman, B., Solorio, T., & Choudhury, M. (2024). The Zeno’s paradox of ‘low-resource’ languages. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 17753–17774). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.983
  13. Office of the Registrar General & Census Commissioner, India. (2011). Census of India 2011. Ministry of Home Affairs, Government of India. https://censusindia.gov.in
  14. Ovadya, A., et al. (2024). Participatory approaches in AI development and governance (arXiv:2407.13100). arXiv.
  15. Research Data Alliance International Indigenous Data Sovereignty Interest Group. (2019, September). CARE Principles for Indigenous Data Governance. The Global Indigenous Data Alliance. https://www.gida-global.org/care
  16. Sloane, M., Moss, E., Awomolo, O., & Forlano, L. (2022). Participation is not a design fix for machine learning. In Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO ’22) (Article 1, pp. 1–6). Association for Computing Machinery. https://doi.org/10.1145/3551624.3555285
  17. Vaughn, L. M., & Jacquez, F. (2020). Participatory research methods: Choice points in the research process. Journal of Participatory Research Methods, 1(1), 13244. https://doi.org/10.35844/001c.13244