The notion of neutral Spanish —also referred to as international, global, or standard Spanish— has long been the subject of debate within Spanish linguistics, particularly in translation and interpreting studies. Early discussions framed neutral Spanish as a pragmatic compromise aimed at maximizing intercomprehension across the Spanish-speaking world, especially in audiovisual translation, journalism, and institutional communication (Montilla Martos 2001; Llorente Pinto 2006). Subsequent scholarship, however, has increasingly problematized this notion, highlighting its ideological aspects, its uneven sociolinguistic consequences, and its tendency to privilege certain regional norms while erasing others (Gutiérrez Maté 2017; Mazzitelli and Garrido Domené 2019; Martínez Moreno 2022). Parallel debates have emerged in the study of languages for specific purposes, where neutralization strategies often conflict with terminological precision and professional conventions (García Izquierdo 2006, 2009), as well as in language policy research, where neutral Spanish has been examined as a tool of symbolic power and linguistic governance in transnational contexts (Paffey 2010, 2014; Amorós Negre 2020; Bravo García 2022).
In recent years, these discussions have acquired renewed urgency due to the rapid advances in artificial intelligence, particularly in neural machine translation and large language models. Much as mass media once contributed to the diffusion and naturalization of certain “neutral” varieties, AI-driven language technologies now play a central role in shaping contemporary norms of written Spanish at a global scale (Muñoz-Basols et al. 2024). Unlike earlier media technologies, however, AI systems operate through opaque processes of data aggregation, annotation —or lack thereof— and statistical optimization, raising new questions about how linguistic diversity is represented, suppressed, or reconfigured in computational environments. Against this backdrop, the present paper examines neutral Spanish not as a sociolinguistic compromise negotiated among speakers, but as a computational artifact produced by specific design choices in AI-assisted translation systems.
The paper argues that the emergence of neutral Spanish in systems such as Google Translate, DeepL, ChatGPT, and Gemini is less the result of an explicit linguistic theory than of the systematic absence of (macro)dialectal annotation and robust data modeling decisions (Gutiérrez Rubio 2024). Drawing on sociolinguistic critiques of standardization and linguistic neutrality, the study analyzes how these platforms generate Spanish outputs that flatten geolectal diversity and implicitly naturalize a small subset of regional features as universally acceptable. Particular attention is devoted to DeepL’s distinction between “Spanish” and “Spanish (Latin America),” a categorization that proves linguistically incoherent when examined against contemporary models of Spanish as a pluricentric language (Lipski 1994; Oesterreicher 2004). Rather than reflecting the internal diversity of American Spanish, this binary collapses multiple macrodialects into a single residual category, thereby reinforcing the idea of a homogeneous Latin American Spanish that lacks empirical grounding.
To address these issues, the study adopts a methodology that integrates annotation structures, artificial intelligence and machine learning analysis, and data modeling. The first methodological component consists of the development of an experimental annotation schema for Spanish macrodialects, grounded in Hispanic dialectology and sociolinguistics. This schema distinguishes among Peninsular, Mexican–Central American, Caribbean, Andean, Chilean, and Rioplatense varieties, drawing on well-established linguistic descriptions rather than market-driven regional labels. The schema is applied to a curated corpus of machine-translated texts produced by the previously mentioned AI-driven platforms. The corpus spans three domains that are particularly relevant to contemporary AI-mediated communication: academic abstracts, institutional and administrative texts, and cultural materials. The annotation process focuses on lexical choices that have been identified in the literature (Moreno Fernández 2020) as salient indicators of macrodialectal variation. By explicitly labeling dialectal features that AI systems leave unmarked, the annotation layer makes visible the patterns of regional dominance and suppression that underpin so-called neutral outputs. Preliminary results reveal a disproportionate presence of Peninsular features in Europe-oriented outputs and of Mexican Spanish features in U.S.-oriented or globally unspecified outputs, while Caribbean, Andean, Chilean, and Rioplatense features are systematically underrepresented or avoided.
The second methodological component situates the annotated corpus within the broader context of neural machine translation and large language models. The paper examines how training on large-scale, heterogeneous, and largely unlabelled data sources —such as Wikipedia, Linguee, and social media platforms— contributes to the emergence of statistically dominant but sociolinguistically opaque varieties of Spanish. While these datasets offer unparalleled volume and coverage, they typically lack reliable metadata regarding regional origin, register, or communicative context. From a model-epistemological perspective, the paper argues that the absence of macrodialect annotation functions as a form of implicit normalization within machine learning pipelines, where the most frequent or institutionally prestigious features come to stand in for the language as a whole. This perspective allows the paper to move beyond critiques that frame AI bias solely in terms of fairness or representation (Samardžić & Ljubešić 2021), instead highlighting the deeper epistemic consequences of unannotated linguistic data.
The third methodological component of the study consists of a data modeling intervention aimed at addressing the limitations identified in current AI systems. Building on the insights gained from annotation and model analysis, the paper proposes a Spanish macrodialect-aware representation model designed for multilingual AI environments. This model departs from existing approaches in two key respects. First, it relies on more reliable and explicitly regionalized lexicographic resources, such as Jergas Hispanas and the Diccionario del Español Mexicano, rather than on undifferentiated web corpora alone. Second, it conceptualizes macrodialects not as mutually exclusive or rigid categories, but as overlapping distributions linked to region, register, and communicative domain.
This representation model reflects contemporary linguistic understandings of variation as gradient and context-dependent, while remaining compatible with computational requirements. By modeling dialectal features as probabilistic tendencies rather than fixed labels, the proposed approach avoids reifying dialect boundaries while still making variation explicit. The paper argues that such a model could be integrated into AI translation systems at multiple levels, from training data selection and annotation to user-facing interface options that go beyond simplistic regional binaries.
References
Amorós Negre, Carla (2020). Los procesos de restandarización lingüística en la hispanofonía: prescripción y norma mediática de la CNN en Español. En Greußlich / F, Lebsanft (eds.), El español, lengua pluricéntrica: discurso, gramática, léxico y medios de comunicación masiva, 271-296. Bonn: Bonn University Press.
Bravo García, Eva (2022). La globalización del español. Estado de la cuestión. Observatorio IEAL sobre América Latina, 3: 1-28.
García Izquierdo, Isabel (2006). El español neutro y la traducción de los lenguajes de especialidad. Sendebar, 17: 149-167.
García Izquierdo, Isabel (2009). El español neutro en los discursos de especialidad: ¿mito, utopía o realidad? Íkal. Revista de lenguaje y cultura, 14(23): 15-39.
Gutiérrez Maté, Miguel (2017). El llamado español latino de los doblajes cinematográficos en la encrucijada entre el español mexicano, el español general y el español neutro. En S. Jansen / G. Müller (eds.), La traducción desde, en y hacia Latinoamérica perspectivas literarias y lingüísticas, 247-274. Madrid/Frankfurt: Iberoamericana/Vervuert.
Gutiérrez Rubio, Enrique (2024). Traducción automática e inteligencia artificial: miradas sobre un fenómeno en vertiginoso avance. Lenguas V;vas, 20: 23-35.
Lipski, John M. (1994). Latin American Spanish. Longman.
Llorente Pinto, María Rosario (2006). ¿Qué es el español neutro?. Cuadernos del Lazarillo, 31: 1-10.
Martínez Moreno, Enrique (2022). El español neutro en el doblaje latino: la imposición a través de luchas simbólicas. Global Media Journal México, 19(36): 1-26.
Mazzitelli, Chiara / Garrido Domené, Fuensanta (2019). Las variedades del español a través del doblaje cinematográfico. Anuario de letras, lingüística y filología, 7(2): 63-82.
Montilla Martos, Antonio (2001). El español «neutro» de los doblajes: intenciones y realidades. Carabela, 50:191-194.
Moreno Fernández, Francisco (2020). Variedades de la lengua española. Londres: Routledge.
Muñoz-Basols, Javier / Palomares Marín, María del Mar / Moreno Fernandez, Francisco (2024). The Digital Linguistic Bias (DLB) in Artificial Intelligence: Implications for Large Language Models in Spanish. Lengua y Sociedad, 23(2): 623-648.
Oesterreicher, Wulf (2004): El español, lengua pluricéntrica. El español en el mundo. Instituto Cervantes.
Paffey, Darren (2010): Globalizing standard Spanish: The promotion of ‘panhispanism’ by Spain’s language guardians. In S. Johnson / T. Milani (eds.), Language Ideologies and Media Discourse: Texts, Practices, Politics, 41-60. London/New York: Continuum.
Paffey, Darren (2014): Language Ideologies and the Globalization of 'Standard' Spanish. London/New York: Bloomsbury.
Samardžić, Tania / Ljubešić, Nikola (2021): Data collection and representation for similar languages, varieties, and dialects. In M. Zampieri & P. Nakov (Eds.), Similar languages, varieties, and dialects: A computational perspective (pp. 121–137). Cambridge University Press.