DH 2026

Daejeon, July 27–31

Thu, July 3013:40–15:10S064106
Short Paper

Sentiments toward migrants and minorities: fine-tuning transformer models for sentiment analysis of historical newspapers

Lucija Brozić (Krušić)
University of Graz, Austria · lucija.krusic@uni-graz.at

This presentation examines how migrants and minorities were represented in Austrian newspaper discourse between 1703 and 1938 by fine-tuning transformer-based models for sentiment analysis. Building on the topic-specific MigraAnno corpus (Brozić, 2025a) and the sentiment-annotated SentiAnno 2.0 corpus (Brozić, 2025b), it focuses on (1) fine-tuning and evaluating historical sentiment models and (2) interpreting long-term sentiment evolution in relation to political orientation on topics of migration and minorities.

The analysis is performed on the MigraAnno corpus, consisting of 96,871 sentences from eight Viennese newspapers representing distinct ideological positions: the monarchist Wienerisches Diarium, conservative Österreichischer Beobachter, liberal Der Wanderer and Neue Freie Presse, Catholic-conservative Das Vaterland, social-democratic Arbeiter-Zeitung, apolitical mass-paper Illustrierte Kronen Zeitung and German-nationalist Deutsches Volksblatt (Paupié, 1960; Melischek & Seethaler, 2016). Sentences were automatically annotated with topics and classified into four analytical categories: MIG, MIN, FAS and CTX, reflecting scholarship on migration, minorities and assimilation in Austria (Brix, 2001; Komlosy, 2004; Hahn, 2000; Steidl, 2020).

Seven German transformer models were fine-tuned for sentiment analysis using the SentiAnno 2.0 gold standard corpus (Brozić, 2025b). These models included contemporary ones, such as GottBERT (Scheible et al., 2024) and bert-base-german-cased (MDZ Digital Library Team, 2019), as well as historical ones, including hmBERT (Schweter et al., 2022) and GHisBERT (Beck & Köllner, 2023), and BERT and ELECTRA models pre-trained on Europeana historical newspapers (Schweter, 2020).

The original SentiAnno annotation scheme included four sentiment labels: “negative” , “neutral”, “positive” and “mixed”.  However, initial experiments showed that including the rarely occurring, conceptually ambiguous “mixed” label substantially reduced model accuracy because the models failed to learn consistent patterns from these cases. Therefore, with the help of annotator comments, the “mixed” annotations were consolidated into the sentiment category they leaned towards, improving model accuracy.

The resulting corpus comprises 2,594 annotated text segments labeled as “negative”, “neutral”, or “positive”, displaying a significant class imbalance: 49.8% are negative, 39.3% are neutral, and 10.9% are positive. The remaining class imbalance was mitigated via class-weighted loss and stratified splits. Despite being trained on contemporary German, GottBERT achieved the best performance (accuracy: 0.74; macro-F1: 0.68), outperforming historically pretrained models. This is likely because GottBERT was pretrained on substantially more data than the historical models (the German portion of the OSCAR dataset, 121 GB; Ortiz Suárez et al., 2019).

The confusion matrix (see Figure 1) of the fine-tuned SentiAnno-GottBERT model shows that negative sentiment is captured reliably, neutral reasonably well, while predicting positive sentiment remains difficult. Sarcasm, irony and ideologically affirmative yet lexically aggressive language (for instance in antisemitic discourse) remain challenging.

Figure 1. GottBERT confusion matrix

Applied to MigraAnno, the model confirms that discourse on migrants and minorities is predominantly negative and neutral (Hampton, 2008; de Veen & Thomas, 2022). Over time, sentiment trajectories correlate with political alignment: low polarity aligns with censorship or conformity, while high polarity marks ideological activism.

In Wienerisches Diarium, MIN topics on Jews, Turks and Croats show strong negative sentiment persisting after the 1781 Edict of Toleration, while FAS topics on multilingualism and German display positive sentiment aligned with Enlightenment ideals (Pesalj, 2019; Scott, 1990; Klingenstein, 1993; Evans, 2004). In the censored early 19th century, Österreichischer Beobachter and Der Wanderer reported negatively on migration, exiles and revolutionary movements (Emerson, 1968; Heyck, 1989). Late 19th-century newspapers diverged: Neue Freie Presse showed strong negative sentiment toward Slavic minorities while expressing negative sentiment toward antisemitism, whereas Catholic Das Vaterland showed negative sentiment toward Jews but more neutral to mildly positive sentiment toward Slavic Catholics. In the early 20th century, Arbeiter-Zeitung combined negative sentiment toward policies driving emigration with positive sentiment toward refugees, Deutsches Volksblatt exhibited persistently negative and exclusionary sentiment, and Illustrierte Kronen Zeitung remained largely neutral before turning increasingly positive toward National Socialism.

A close reading revealed a core limitation of sentence-level polarity. “Negative” sentiment sometimes reflected attitudes toward government policy while implicitly expressing empathy toward migrants or minority groups. For instance, the negativity expressed by Arbeiter-Zeitung regarding "Galicia" was directed at governmental shortcomings and demonstrated solidarity with refugees. In contrast, the negativity exhibited by Deutsches Volksblatt was directed at refugees. Similarly, the negativity displayed by Neue Freie Presse regarding the topic “Jews” often signified criticism of antisemitism. Aspect-based sentiment analysis and stance detection may offer more fine-grained alternatives in the future (Dejaeghere et al., 2024; Hamdi et al., 2021). Improvements to model accuracy could include domain-adaptive pretraining on historical corpora, expanding positive and figurative examples, and reviewing borderline cases with experts, as in active learning (Brunner et al., 2020; Schmidt et al., 2021; Prabhu et al., 2021).

This presentation contributes to the growing body of work on sentiment analysis in the Digital Humanities and presents novel fine-tuned models for sentiment analysis of historical texts in German. The best-performing model, SentiAnno-GottBERT (Brozić, 2025c), is openly available on HuggingFace, while the corpora are available on Zenodo, following FAIR principles (Wilkinson et al., 2016). Together, these findings underscore the analytical potential and limits of sentiment analysis in historical newspaper research, highlighting the necessity of integrating computational methods with contextual and source-critical interpretations.

References
  1. Beck, Christin / Köllner, Marisa (2023): “GHisBERT – Training BERT from scratch for lexical semantic investigations across historical German language stages”, in: Proceedings of the 4th Workshop on Computational Approaches to Historical Language Change, 33–45. Singapore: Association for Computational Linguistics.
  2. Becker, Peter (2010): “Governance of Migration in the Habsburg Empire and the Republic of Austria”, in: National Approaches to the Administration of International Migration, hrsg. von Peri E. Arnold, 32–52. Amsterdam: IOS Press.
  3. Brix, Emil (2001): “Assimilation in the Late Habsburg Monarchy”, in: Österreich-Konzeptionen und jüdisches Selbstverständnis, 29–42. Tübingen: Max Niemeyer Verlag.
  4. Brozić, Lucija (2025a): lukru/SentiAnno-GottBERT. Hugging Face.
  5. Brozić, Lucija (2025b): MigraAnno. Zenodo.
  6. Brozić, Lucija (2025c): SentiAnno 2.0. Zenodo.
  7. Brunner, Annelen / Tu, Ngoc Duyen Tanja / Weimer, Lukas / Jannidis, Fotis (2020): “To BERT or not to BERT – Comparing Contextual Embeddings in a Deep Learning Architecture for the Automatic Recognition of four Types of Speech, Thought and Writing Representation”, in: Proceedings of the 5th Swiss Text Analytics Conference (SwissText) & 16th Conference on Natural Language Processing (KONVENS), hrsg. von Sarah Ebling, Don Tuggener, Mark Cieliebak und Martin Volk.
  8. De Veen, Linda / Thomas, Richard (2022): “Shooting for neutrality? Analysing bias in terrorism reports in Dutch newspapers”, in: Media, War & Conflict 15, 2: 146–164.
  9. Dejaeghere, Tess / Singh, Pranaydeep / Lefever, Els / Birkholz, Julie M. (2024): “Exploring aspect-based sentiment analysis methodologies for literary-historical research purposes”, in: Proceedings of the Third Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC-COLING-2024, 129–143.
  10. Emerson, Donald E. (1968): “Subversion in Austria and Germany”, in: Metternich and the Political Police: Security and Subversion in the Hapsburg Monarchy (1815–1830), hrsg. von Donald E. Emerson, 100–135. Dordrecht: Springer Netherlands.
  11. Evans, R. J. W. (2004): “Language and State Building: The Case of the Habsburg Monarchy”, in: Austrian History Yearbook 35: 1–24.
  12. Hahn, Sylvia (2000): “Inclusion and Exclusion of Migrants in the Multicultural Realm of the Habsburg ‘State of Many Peoples’”, in: Histoire sociale / Social History.
  13. Hamdi, Ahmed / Pontes, Elvys Linhares / Boros, Emanuela / Nguyen, Thi Tuyet Hai / Hackl, Günter / Moreno, Jose G. / Doucet, Antoine (2021): “A Multilingual Dataset for Named Entity Recognition, Entity Linking and Stance Detection in Historical Newspapers”, in: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2328–2334. Virtual Event Canada: ACM.
  14. Hampton, Mark (2008): “The ‘Objectivity’ Ideal and Its Limitations in 20th-Century British Journalism”, in: Journalism Studies 9, 4: 477–493.
  15. Heyck, Eduard (1989): Die Allgemeine Zeitung, 1798–1898: Beiträge zur Geschichte der deutschen Presse. Nabu Press.
  16. Klingenstein, Grete (1993): “Modes of Religious Tolerance and Intolerance in Eighteenth-Century Habsburg Politics”, in: Austrian History Yearbook 24: 1–16.
  17. Komlosy, Andrea (2004): “State, Regions, and Borders: Single Market Formation and Labor Migration in the Habsburg Monarchy, 1750–1918”, in: Review (Fernand Braudel Center) 27, 2: 135–177.
  18. MDZ Digital Library team (dbmdz) (2019): bert-base-german-cased. Hugging Face.
  19. Melischek, Gabriele / Seethaler, Josef (2016): “Die Tagespresse der franzisko-josephinischen Ära”, in: Österreichische Mediengeschichte, hrsg. von Matthias Karmasin und Christian Oggolder, 167–192. Wiesbaden: Springer Fachmedien Wiesbaden.
  20. Ortiz Suárez, Pedro Javier / Sagot, Benoît / Romary, Laurent (2019): “Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures”, in: Proceedings of the Workshop on Challenges in the Management of Large Corpora, 337 KB.
  21. Paupié, Kurt (1960): Handbuch der österreichischen Pressegeschichte, 1848–1959: Die zentralen pressepolitischen Einrichtungen des Staates. Wien: W. Braumüller.
  22. Pesalj, J. (2019): Monitoring migrations: the Habsburg-Ottoman border in the eighteenth century. Dissertation, Leiden: Leiden University.
  23. Prabhu, Sumanth / Mohamed, Moosa / Misra, Hemant (2021): “Multi-class Text Classification using BERT-based Active Learning”, in: arXiv.
  24. Schmidt, Thomas / Dennerlein, Katrin / Wolff, Christian (2021): “Using Deep Learning for Emotion Analysis of 18th and 19th Century German Plays”, in: Fabrikation von Erkenntnis: Experimente in den Digital Humanities.
  25. Schweter, Stefan (2020a): Europeana BERT and ELECTRA models. Zenodo.
  26. --- (2020b): Europeana BERT and ELECTRA models. Zenodo.
  27. Schweter, Stefan / März, Luisa / Schmid, Katharina / Çano, Erion (2022): “hmBERT: Historical Multilingual Language Models for Named Entity Recognition”, in: arXiv.
  28. Steidl, Annemarie (2020a): On Many Routes: Internal, European, and Transatlantic Migration in the Late Habsburg Empire. Central European Studies.
  29. --- (2020b): On Many Routes: Internal, European, and Transatlantic Migration in the Late Habsburg Empire. Central European Studies.
  30. Wilkinson, Mark D. / Dumontier, Michel / Aalbersberg, IJsbrand Jan / Appleton, Gabrielle / Axton, Myles / Baak, Arie / Blomberg, Niklas / u. a. (2016): “The FAIR Guiding Principles for scientific data management and stewardship”, in: Scientific Data 3, 1: 160018.