Daejeon, July 27–31
Language, being a powerful tool, can be manipulated by malicious actors to shape public opinion. In this study, we examine how pro-Kremlin propaganda about the Russo-Ukrainian War is linguistically framed in a Russian Wikipedia Fork (RWFork) compared to the original version. Propaganda studies in the context of this war have mostly focused on media analysis (Akhynko et al. 2025, Hein 2023, Solopova et al. 2023, Vanetik et al. 2023). (Akhynko et al. 2025, Hein 2023, Solopova et al. 2023, Vanetik et al. 2023). Alyukov et al. (2025) examined the differences between state-controlled traditional media (such as press and TV) and social media, showing that the former rely on demobilization by normalizing the war, while the latter have a more mobilizational character.Alyukov et al. (2025) examined the differences between state-controlled traditional media (such as press and TV) and social media, showing that the former rely on demobilization by normalizing the war, while the latter have a more mobilizational character.
Wikipedia presents another text type: despite aiming to be neutral and objective, it can also enable knowledge manipulation, specifically in its alternative versions such as RWFork (Trokhymovych et al. 2025). Created in June 2023, RWFork is a Russia-law-compliant copy of the Russian Wikipedia (Trokhymovych et al. 2025). Created in June 2023, RWFork is a Russia-law-compliant copy of the Russian Wikipedia (RW; Cohen 2023). (RW; Cohen 2023). Trokhymovych et al. (2025) have shown that most of the edits in RWFork are related to Russia’s full-scale invasion of Ukraine. These changes involve territorial claims, war terminology, and sanctions-related information.Trokhymovych et al. (2025) have shown that most of the edits in RWFork are related to Russia’s full-scale invasion of Ukraine. These changes involve territorial claims, war terminology, and sanctions-related information.
Our study uses the RWFork dataset (Trokhymovych et al. 2025), containing edits from 1.9M page titles between May and September 2023. We analyze edits involving knowledge manipulation related to the 2022 Russian invasion of Ukraine by selecting articles labeled as Territorial Claims Dispute, Terminology Changes Ukraine, and Sanctions Edit Adjustments in the original corpus. The filtered dataset contains 13,048 articles and 92,629 sentences (reduced from 100,738 after deduplication).(Trokhymovych et al. 2025), containing edits from 1.9M page titles between May and September 2023. We analyze edits involving knowledge manipulation related to the 2022 Russian invasion of Ukraine by selecting articles labeled as Territorial Claims Dispute, Terminology Changes Ukraine, and Sanctions Edit Adjustments in the original corpus. The filtered dataset contains 13,048 articles and 92,629 sentences (reduced from 100,738 after deduplication).
We employ Kullback-Leibler Divergence (KLD; Kullback / Leibler 1951), which offers an interpretable way to identify distinctive linguistic features. Unlike other keyness metrics (cf. (KLD; Kullback / Leibler 1951), which offers an interpretable way to identify distinctive linguistic features. Unlike other keyness metrics (cf. Du et al. 2022), KLD has the advantage of considering over the whole frequency band and has been widely used in studies of language variation and change Du et al. 2022), KLD has the advantage of considering over the whole frequency band and has been widely used in studies of language variation and change (Bizzoni et al. 2020, Bochkarev et al. 2014, Degaetano-Ortlieb / Teich 2022, Hughes et al. 2012, Klingenstein et al. 2014). KLD is used to measure how much the language of RWFork diverges from RW by quantifying the extra information needed to represent one probability distribution with the other, thereby identifying key linguistic features (e.g., words) that contribute to the linguistic differences. For instance, to calculate the KLD of RW’s language, given that of RWFork, we use this formula:(Bizzoni et al. 2020, Bochkarev et al. 2014, Degaetano-Ortlieb / Teich 2022, Hughes et al. 2012, Klingenstein et al. 2014). KLD is used to measure how much the language of RWFork diverges from RW by quantifying the extra information needed to represent one probability distribution with the other, thereby identifying key linguistic features (e.g., words) that contribute to the linguistic differences. For instance, to calculate the KLD of RW’s language, given that of RWFork, we use this formula:
We focused only on noun, proper noun, verb, adjective, and adverb lemmas, which are most likely to reflect meaningful content differences between the two Wikipedia versions.
Our analysis of the 50 most distinctive words for RW and RWFork revealed many parallels with prior research on the divergent language of the media reporting on the Russo-Ukrainian War (see Table 1; higher KLD values indicate higher distinctiveness). For example, the terms invasion, war, and aggression are among the most prominent KLD words for the original Wikipedia version; they are either removed from RWFork or substituted with vague expressions such as military actions and special military operation to downplay the invasion. Similar euphemistic substitutions appear in media propaganda, specifically in state-affiliated outlets as opposed to independent ones (Park et al. 2022), pro-Russian vs. pro-Ukrainian Telegram channels (Park et al. 2022), pro-Russian vs. pro-Ukrainian Telegram channels (Ustyianovych / Barbosa 2024), and traditional media vs. social media (Ustyianovych / Barbosa 2024), and traditional media vs. social media (Alyukov et al. 2022, Vestel / Degaetano-Ortlieb 2025). In addition, words like (Alyukov et al. 2022, Vestel / Degaetano-Ortlieb 2025). In addition, words like occupation and annexation are replaced with euphemisms such as inclusion. Moreover, the recognition of Ukraine’s statehood in RW is absent from RWFork, with words like independence and sovereignty being removed.
Furthermore, about one-third of the top KLD words for RWFork are names of certain Ukrainian territories either occupied or claimed by Russia. Notably, RWFork clearly recognizes the self-proclaimed Russia-backed DPR (Donetsk People’s Republic) and LPR (Luhansk People’s Republic), since their names are among the most distinctive words for this version. These and similar changes demonstrate the Kremlin’s aim to convince the population that parts of Ukraine have historically belonged to Russia, pointing to a territorial control propaganda, which also reflects the normalization frame of the demobilizational approach characteristic of traditional media (Vestel / Degaetano-Ortlieb 2025).(Vestel / Degaetano-Ortlieb 2025).
Table 1. Examples of the most distinctive words across Wikipedia versions.
Our analysis demonstrates that the language of RWFork aligns closely with pro-Kremlin propaganda spread in pro-Russian, state-affiliated, or traditional media, as opposed to the original version. This is evidenced by the efforts to downplay the war and the focus on geopolitical control, which point to a demobilizational strategy. Overall, our work offers insights into how propaganda is constructed in political environments and demonstrates the potential of reproducible computational techniques for its detection.