DH 2026

Daejeon, July 27–31

Poster

Rewriting Reality: Tracking Russian Propaganda Through Wikipedia

Anastasiia Vestel
Saarland University, Germany · anastasiia.vestel@uni-saarland.de
Stefania Degaetano-Ortlieb
Saarland University, Germany · s.degaetano@mx.uni-saarland.de

Introduction

Language, being a powerful tool, can be manipulated by malicious actors to shape public opinion. In this study, we examine how pro-Kremlin propaganda about the Russo-Ukrainian War is linguistically framed in a Russian Wikipedia Fork (RWFork) compared to the original version. Propaganda studies in the context of this war have mostly focused on media analysis (Akhynko et al. 2025, Hein 2023, Solopova et al. 2023, Vanetik et al. 2023). (Akhynko et al. 2025, Hein 2023, Solopova et al. 2023, Vanetik et al. 2023). Alyukov et al. (2025) examined the differences between state-controlled traditional media (such as press and TV) and social media, showing that the former rely on demobilization by normalizing the war, while the latter have a more mobilizational character.Alyukov et al. (2025) examined the differences between state-controlled traditional media (such as press and TV) and social media, showing that the former rely on demobilization by normalizing the war, while the latter have a more mobilizational character.

Wikipedia presents another text type: despite aiming to be neutral and objective, it can also enable knowledge manipulation, specifically in its alternative versions such as RWFork (Trokhymovych et al. 2025). Created in June 2023, RWFork is a Russia-law-compliant copy of the Russian Wikipedia (Trokhymovych et al. 2025). Created in June 2023, RWFork is a Russia-law-compliant copy of the Russian Wikipedia (RW; Cohen 2023). (RW; Cohen 2023). Trokhymovych et al. (2025) have shown that most of the edits in RWFork are related to Russia’s full-scale invasion of Ukraine. These changes involve territorial claims, war terminology, and sanctions-related information.Trokhymovych et al. (2025) have shown that most of the edits in RWFork are related to Russia’s full-scale invasion of Ukraine. These changes involve territorial claims, war terminology, and sanctions-related information.

Methodology

Our study uses the RWFork dataset (Trokhymovych et al. 2025), containing edits from 1.9M page titles between May and September 2023. We analyze edits involving knowledge manipulation related to the 2022 Russian invasion of Ukraine by selecting articles labeled as Territorial Claims Dispute, Terminology Changes Ukraine, and Sanctions Edit Adjustments in the original corpus. The filtered dataset contains 13,048 articles and 92,629 sentences (reduced from 100,738 after deduplication).(Trokhymovych et al. 2025), containing edits from 1.9M page titles between May and September 2023. We analyze edits involving knowledge manipulation related to the 2022 Russian invasion of Ukraine by selecting articles labeled as Territorial Claims Dispute, Terminology Changes Ukraine, and Sanctions Edit Adjustments in the original corpus. The filtered dataset contains 13,048 articles and 92,629 sentences (reduced from 100,738 after deduplication).

We employ Kullback-Leibler Divergence (KLD; Kullback / Leibler 1951), which offers an interpretable way to identify distinctive linguistic features. Unlike other keyness metrics (cf. (KLD; Kullback / Leibler 1951), which offers an interpretable way to identify distinctive linguistic features. Unlike other keyness metrics (cf. Du et al. 2022), KLD has the advantage of considering over the whole frequency band and has been widely used in studies of language variation and change Du et al. 2022), KLD has the advantage of considering over the whole frequency band and has been widely used in studies of language variation and change (Bizzoni et al. 2020, Bochkarev et al. 2014, Degaetano-Ortlieb / Teich 2022, Hughes et al. 2012, Klingenstein et al. 2014). KLD is used to measure how much the language of RWFork diverges from RW by quantifying the extra information needed to represent one probability distribution with the other, thereby identifying key linguistic features (e.g., words) that contribute to the linguistic differences. For instance, to calculate the KLD of RW’s language, given that of RWFork, we use this formula:(Bizzoni et al. 2020, Bochkarev et al. 2014, Degaetano-Ortlieb / Teich 2022, Hughes et al. 2012, Klingenstein et al. 2014). KLD is used to measure how much the language of RWFork diverges from RW by quantifying the extra information needed to represent one probability distribution with the other, thereby identifying key linguistic features (e.g., words) that contribute to the linguistic differences. For instance, to calculate the KLD of RW’s language, given that of RWFork, we use this formula:

We focused only on noun, proper noun, verb, adjective, and adverb lemmas, which are most likely to reflect meaningful content differences between the two Wikipedia versions.

Results

Our analysis of the 50 most distinctive words for RW and RWFork revealed many parallels with prior research on the divergent language of the media reporting on the Russo-Ukrainian War (see Table 1; higher KLD values indicate higher distinctiveness). For example, the terms invasion, war, and aggression are among the most prominent KLD words for the original Wikipedia version; they are either removed from RWFork or substituted with vague expressions such as military actions and special military operation to downplay the invasion. Similar euphemistic substitutions appear in media propaganda, specifically in state-affiliated outlets as opposed to independent ones (Park et al. 2022), pro-Russian vs. pro-Ukrainian Telegram channels (Park et al. 2022), pro-Russian vs. pro-Ukrainian Telegram channels (Ustyianovych / Barbosa 2024), and traditional media vs. social media (Ustyianovych / Barbosa 2024), and traditional media vs. social media (Alyukov et al. 2022, Vestel / Degaetano-Ortlieb 2025). In addition, words like (Alyukov et al. 2022, Vestel / Degaetano-Ortlieb 2025). In addition, words like occupation and annexation are replaced with euphemisms such as inclusion. Moreover, the recognition of Ukraine’s statehood in RW is absent from RWFork, with words like independence and sovereignty being removed.

Furthermore, about one-third of the top KLD words for RWFork are names of certain Ukrainian territories either occupied or claimed by Russia. Notably, RWFork clearly recognizes the self-proclaimed Russia-backed DPR (Donetsk People’s Republic) and LPR (Luhansk People’s Republic), since their names are among the most distinctive words for this version. These and similar changes demonstrate the Kremlin’s aim to convince the population that parts of Ukraine have historically belonged to Russia, pointing to a territorial control propaganda, which also reflects the normalization frame of the demobilizational approach characteristic of traditional media (Vestel / Degaetano-Ortlieb 2025).(Vestel / Degaetano-Ortlieb 2025).

Table 1. Examples of the most distinctive words across Wikipedia versions.

Conclusion

Our analysis demonstrates that the language of RWFork aligns closely with pro-Kremlin propaganda spread in pro-Russian, state-affiliated, or traditional media, as opposed to the original version. This is evidenced by the efforts to downplay the war and the focus on geopolitical control, which point to a demobilizational strategy. Overall, our work offers insights into how propaganda is constructed in political environments and demonstrates the potential of reproducible computational techniques for its detection.

References
  1. Akhynko, Kateryna / Kosovan, Oleksandr / Trokhymovych, Mykola (2025): “Hidden Persuasion: Detecting Manipulative Narratives on Social Media During the 2022 Russian Invasion of Ukraine”, in: Romanyshyn, Mariana (ed.): Proceedings of the Fourth Ukrainian Natural Language Processing Workshop (UNLP 2025). Vienna, Austria (online): Association for Computational Linguistics 194–202. DOI: 10.18653/v1/2025.unlp-1.19.
  2. Alyukov, Maxim / Kunilovskaya, Maria / Semenov, Andrei (2025): “Confuse and Normalise: Authoritarian Propaganda in a High-Choice Media Environment and Russia's Invasion of Ukraine”, in: Goode, Paul (ed.): Russian Propaganda Today: Challenges, Effectiveness, and Resistance. University of Michigan Press, University of Manchester Press: in print.
  3. Alyukov, Maxim / Semenov, Andrei / Kunilovskaya, Maria (2022): “Propaganda Setbacks and Appropriation of Anti-War Language: ‘Special Military Operation’ in Russian Mass Media and Social Networks (February–July 2022). Monitoring Report №1” <https://research.manchester.ac.uk/en/publications/propaganda-setbacks-and-appropriation-of-anti-war-language-specia/> [23.04.2026].
  4. Bizzoni, Yuri / Degaetano-Ortlieb, Stefania / Fankhauser, Peter / Teich, Elke (2020): “Linguistic Variation and Change in 250 Years of English Scientific Writing: A Data-Driven Approach”, in: Frontiers in Artificial Intelligence 3: 73. DOI: 10.3389/frai.2020.00073.
  5. Bochkarev, Vladimir / Solovyev, Valery / Wichmann, Søren (2014): “Universals versus Historical Contingencies in Lexical Evolution”, in: Journal of The Royal Society Interface 11, 101: 20140841. DOI: 10.1098/rsif.2014.0841.
  6. Cohen, Noam (2023): “Russian Wikipedia’s Top Editor Leaves to Launch a Putin-Friendly Clone”, in: Bloomberg.com <https://www.bloomberg.com/news/articles/2023-07-12/russian-wikipedia-editor-leaves-to-launch-a-putin-friendly-clone> [23.04.2026].
  7. Degaetano-Ortlieb, Stefania / Teich, Elke (2022): “Toward an Optimal Code for Communication: The Case of Scientific English”, in: Corpus Linguistics and Linguistic Theory 18, 1: 175–207. DOI: 10.1515/cllt-2018-0088.
  8. Du, Keli / Dudar, Julia / Schöch, Christof (2022): “Evaluation of Measures of Distinctiveness. Classification of Literary Texts on the Basis of Distinctive Words”, in: Journal of Computational Literary Studies 1, 1. DOI: 10.48694/jcls.102.
  9. Hein, Vitalij (2023): Propaganda Detection in Russian and American News Coverage about the War in Ukraine through Text Classification. Thesis, Technische Universität Wien. DOI: 10.34726/hss.2023.104640.
  10. Hughes, James M. / Foti, Nicholas J. / Krakauer, David C. / Rockmore, Daniel N. (2012): “Quantitative Patterns of Stylistic Influence in the Evolution of Literature”, in: Proceedings of the National Academy of Sciences 109, 20: 7682–7686. DOI: 10.1073/pnas.1115407109.
  11. Klingenstein, Sara / Hitchcock, Tim / DeDeo, Simon (2014): “The Civilizing Process in London’s Old Bailey”, in: Proceedings of the National Academy of Sciences of the United States of America 111, 26: 9419–9424. DOI: 10.1073/pnas.1405984111.
  12. Kullback, Solomon / Leibler, Richard A. (1951): “On Information and Sufficiency”, in: The Annals of Mathematical Statistics 22, 1: 79–86. DOI: 10.1214/aoms/1177729694.
  13. Park, Chan Young / Mendelsohn, Julia / Field, Anjalie / Tsvetkov, Yulia (2022): “Challenges and Opportunities in Information Manipulation Detection: An Examination of Wartime Russian Media”, in: Goldberg, Yoav / Kozareva, Zornitsa / Zhang, Yue (eds.): Findings of the Association for Computational Linguistics: EMNLP 2022, Abu Dhabi, UAE: 5209–5235. DOI: 10.18653/v1/2022.findings-emnlp.382 <https://aclanthology.org/2022.findings-emnlp.382/> [23.04.2026].
  14. Solopova, Veronika / Benzmüller, Christoph / Landgraf, Tim (2023): “The Evolution of Pro-Kremlin Propaganda from a Machine Learning and Linguistics Perspective”, in: Romanyshyn, Mariana (ed.): Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP), Dubrovnik, Croatia: 40–48. DOI: 10.18653/v1/2023.unlp-1.5.
  15. Trokhymovych, Mykola / Kosovan, Oleksandr / Forrester, Nathan / Aragón, Pablo / Saez-Trumper, Diego / Baeza-Yates, Ricardo (2025): “Characterizing Knowledge Manipulation in a Russian Wikipedia Fork”, in: Proceedings of the International AAAI Conference on Web and Social Media 19: 1924–1936. DOI: 10.1609/icwsm.v19i1.35910.
  16. Ustyianovych, Taras / Barbosa, Denilson (2024): “Instant Messaging Platforms News Multi-Task Classification for Stance, Sentiment, and Discrimination Detection”, in: Romanyshyn, Mariana / Romanyshyn, Nataliia / Hlybovets, Andrii / Ignatenko, Oleksii (eds.): Proceedings of the Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024, Torino, Italia: 30–40.
  17. Vanetik, Natalia / Litvak, Marina / Reviakin, Egor / Tyamanova, Margarita (2023): “Propaganda Detection in Russian Telegram Posts in the Scope of the Russian Invasion of Ukraine”, in: Proceedings of the Conference Recent Advances in Natural Language Processing – Large Language Models for Natural Language Processing. INCOMA Ltd., Shoumen, BULGARIA. 1162–1170. DOI: 10.26615/978-954-452-092-2_123.
  18. Vestel, Anastasiia / Degaetano-Ortlieb, Stefania (2025): “From War to Special Military Operation: Interpretable Detection of Linguistic Propaganda Framing in Russian Media”, in: Workshop Proceedings of the 19th International AAAI Conference on Web and Social Media 2025: 50. DOI: 10.36190/2025.50