DH 2026

Daejeon, July 27–31

Wed, July 2914:00–15:30S019105
Long Paper

When Geometry Meets Etymology: Cultural Patterns and Memory in the Language Revealed by PCA and Automatic Classification

Jacek Bąkowski
Institute of Polish Language, Polish Academy of Sciences, Poland · jacek.bakowski@ijp.pan.pl

Abstract

Recent advances in computational linguistics have shown that distributional word embeddings capture semantic, syntactic, and stylistic properties of lexical items (Baroni / Lenci 2010). However, it remains unclear whether these representations also encode historically grounded linguistic information, particularly etymological structure. Although embeddings are trained on contemporary usage, they may still reflect traces of historical language contact and lexical stratification.

This question is especially relevant in languages shaped by extensive borrowing. Hindi offers a useful case study because its lexicon reflects centuries of interaction with Sanskrit and Perso-Arabic sources. Therefore, many semantic domains therefore contain near-synonymous pairs in which one item derives from Sanskrit and the other from Perso-Arabic. While these pairs often share core meaning, they may differ in register, sociolinguistic context, and cultural associations.

Despite substantial research on lexical borrowing in South Asian linguistics, little computational work has examined whether etymological origin leaves detectable traces in modern distributional spaces. This study addresses that gap by testing whether word embeddings encode signals associated with etymological origin in Hindi. Using a dataset of near-synonymous Sanskrit- and Perso-Arabic-derived pairs, it evaluates whether classification models can predict etymological origin from embedding vectors, and whether geometric analysis reveals clustering consistent with historical lexical stratification. At the same time, embeddings may primarily encode sociolinguistic features such as register and style, for which etymology often functions as a proxy.

Historical and linguistic context

India has long been shaped by transregional exchange through the Indian Ocean and the Iranian Plateau. This interconnectedness intensified between the 11th and 18th centuries, when Sanskrit and Persian cultural spheres came into close contact. During this period, Persianate culture spread widely across the subcontinent, becoming both a practical requirement for administrative and literary life and a marker of elite identity. As a result, India emerged as a major center of the Persian-speaking world for several centuries (Eaton 2019: 8, 380–2; Kuczkiewicz-Fraś 2012: 47-51).

Centuries of contact with Persian left a lasting imprint on the languages of South Asia, especially the northern and western vernaculars, including Modern Hindi. Even after Persian declined in the 18th century, Perso-Arabic vocabulary remained deeply integrated into Indian languages, often enriching rather than replacing native elements. Today, Sanskrit-derived and Perso-Arabic words coexist in Modern Hindi in a lexical continuum shaped largely by stylistic and cultural associations, with speakers selecting between them depending on context and register (Kuczkiewicz-Fraś 2003: 10; 2012: 56).

Research Goals and Experimental Design

This study examines whether the historical distinction between Sanskrit-derived and Perso-Arabic vocabulary in Modern Hindi is reflected in the geometric structure of distributional word embeddings. More specifically, it tests whether etymological origin can be inferred from embedding representations, which, by encoding statistical regularities of language, may also capture underlying patterns of lexical structure beyond surface-level meaning, even when lexical items function as near-synonyms in contemporary usage.

To address this question, parallel sets of 135 Sanskrit-derived and Perso-Arabic near-synonyms were selected using McGregor (1993), and their distributional representations were extracted from two independently trained Word2Vec models. The vectors were then input to four machine-learning classifiers—K-Nearest Neighbors, Random Forest, Support Vector Machines, and Naive Bayes—to evaluate whether etymological origin can be reliably predicted from embedding representations.

Global classifier performance and cross-corpus generalizability were assessed through 100 independent classification runs conducted on both a main and a control corpus. In each iteration, synonym pairs were randomly split into training and test sets (80:20), classifiers were trained on embeddings from each corpus, accuracy was evaluated on the held-out data, and misclassification events were recorded for each word.

Potential data leakage through synonym pairs was further controlled by comparing two train–test splitting strategies: a standard word-level split, in which individual words were assigned randomly to train or test sets, and a stricter meaning-level split, in which both members of each synonym pair were assigned exclusively to one set. Finally, a label permutation test was conducted by randomly shuffling class labels in the training data while keeping the input representations unchanged, in order to verify that classification performance was not driven by artefacts or spurious correlations.

To assess the extent to which the results arise from the data itself rather than from a particular classification method, and to confront the supervised classifiers with an unsupervised analysis, PCA was used to examine the structure of the embeddings based solely on their dimensions, without any knowledge of the classes (i.e., origin); only afterward colors were applied to the projection to indicate word origin.

Results: Classifier Performance

All classifiers achieved at least 83% accuracy under both splitting strategies, with consistently low standard deviations across runs (SD = 0.030–0.051; see fig. 1), indicating that the results are not driven by random splits, but reflect stable distributional structure associated with etymological strata.

Figure Classifier performance across methods, corpora and with split on the word and meaning levels. For clarity, only the first 100 iterations are displayed.

In the label permutation test, all classifiers dropped to chance-level performance (~50%) across both corpora and splitting strategies (fig. 2), confirming that the observed accuracy reflects genuine structure in the embedding space, rather than artefacts of the experimental setup.

Figure Classifier performance under label permutation test. The training labels were randomly permuted while feature representations were preserved, yielding chance-level performance across all classifiers and datasets.

Results: PCA Visualizations

The two-dimensional PCA projection revealed the existence of semi-disjoint origin islands in the latent space, with Sanskrit-origin and Perso-Arabic origin words forming partially separated clusters. Interestingly, when connecting individual synonym pairs, words are often closer to others sharing the same etymology, than to their semantic counterpart (fig. 3). This separation became even more pronounced in the three-dimensional PCA projection (fig. 4).

Figure PCA of the synonym embeddings in both corpora. Pairs of synonyms are connected.

Figure 3D PCA of the synonym embeddings in both corpora.

Mapping cumulative classifier runs onto the PCA projection revealed distinct error distributions across classifiers (fig. 5). While all achieve similar aggregate performance, they differed in their treatment of etymologically ambiguous vocabulary: distance-based methods (k-NN) struggle with a broader set of words showing weak local clustering, whereas margin-based approaches (SVM) confined errors to a narrow set of boundary cases in overlapping regions between etymological clusters. Random Forest exhibited the most uniform error distribution, spreading uncertainty more evenly, rather than concentrating on boundary cases. Notably, over 50% of the vocabulary was always correctly classified, indicating clear etymological signals, and these patterns were confirmed on the control corpus.

Figure PCA of the synonym embeddings with classification errors on main corpus.

Conclusions

The study yields several key findings across multiple dimensions. The results are consistent across both corpora, indicating that the observed effects are not marginal but reflect a robust and systematic pattern, suggesting generalizability beyond dataset-specific properties.

The results suggest that etymology is reflected in lexical behavior centuries after borrowing, leaving lasting semantic footprints in contemporary language. However, the observed separability may reflect not only etymological origin itself, but also sociolinguistic properties such as register and stylistic variation, which are historically associated with different lexical strata. Over time, despite linguistic change, these distinctions may remain encoded in distributional patterns, providing computational evidence that language contact can leave persistent, though not necessarily consciously accessible, traces in usage.

PCA visualizations show that words sharing etymology cluster tightly in the latent space, with each stratum forming a coherent group. Members of synonymous pairs are closer to other words of the same origin than to their semantic counterparts. This may suggest the existence of broader patterns in lexical organization: an emergent semantic frame of etymologically related words, which, despite differences in meaning, form distinct networks and conceptual groupings occupying related niches shaped by shared history.

These results also challenge traditional notions of synonymy. The examined lexemes are considered as synonyms, and thus, theoretically interchangeable. However, speakers seem to select a distinct variant depending on context, suggesting that denotational equivalents display systematic distributional differences. This sheds new light on the principle of no synonymy (Goldberg 1995: 67), recently reformulated as the principle of no equivalence (Leclercq & Morin 2023), which holds that any difference in form entails a difference in meaning. It has already been stated that, while two words may be distributionally and referentially equivalent, they can still be associated with distinct prototypes or distinguished with respect to the domains against which they are understood (Taylor 2003: 59, 89). This aligns with the intuition that “even if near-synonyms name one and the same thing, they name it in different ways: they present different perspectives on a situation” (Divjak 2010: 1). This study suggests, however, that such prototyping and perspectival differences may be partly shaped by words’ origin.

These findings provide a foundation for applying the same methodology to other languages with layered lexical histories such as English, where Old English and Norman-French strata coexist, or Polish, where Latin and Gallic borrowings interact with the native lexicon.

Overall, the results highlight both the potential of distributional semantic models to capture historically grounded structure in lexical systems and the existence of layered organization of language, in which different historical strata of the lexicon exhibit distinct distributional behaviors.

References
  1. Baroni M., Lenci, A. (2010). “Distributional memory: A general framework for corpus-based semantics” in: Computational Linguistics, 36(4), 673–721.
  2. Divjak, Dagmar (2010): Structuring the lexicon: A clustered model for near-synonymy. Vol. 43. Walter de Gruyter.
  3. Eaton Richard M. (2019): India in the Persianate Age: 1000–1765. Penguin UK.
  4. Goldberg, Adele (2019): Explain Me This: Creativity, Competition, and the Partial Productivity of Constructions. Princeton: Princeton University Press.
  5. Kuczkiewicz-Fraś, A. (2003): Perso-Arabic Hybrids in Hindi. The Socio-Linguistic and Structural Analysis. New Delhi: Manohar Publishers.
  6. Kuczkiewicz-Fraś, A. (2012): Perso-Arabic Loanwords in Hindustani. Part II, Linguistic Study. Kraków: Księgarnia Akademicka.
  7. Leclercq, Benoît / Morin, Cameron (2023): “No equivalence: A new principle of no synonymy” in: Constructions, 15,1. DOI: 10.24338/cons-535.
  8. McGregor, Ronald S. (1993): The Oxford Hindi-English Dictionary. Oxford University Press.