Daejeon, July 27–31
This presentation examines how migrants and minorities were represented in Austrian newspaper discourse between 1703 and 1938 by fine-tuning transformer-based models for sentiment analysis. Building on the topic-specific MigraAnno corpus (Brozić, 2025a) and the sentiment-annotated SentiAnno 2.0 corpus (Brozić, 2025b), it focuses on (1) fine-tuning and evaluating historical sentiment models and (2) interpreting long-term sentiment evolution in relation to political orientation on topics of migration and minorities.
The analysis is performed on the MigraAnno corpus, consisting of 96,871 sentences from eight Viennese newspapers representing distinct ideological positions: the monarchist Wienerisches Diarium, conservative Österreichischer Beobachter, liberal Der Wanderer and Neue Freie Presse, Catholic-conservative Das Vaterland, social-democratic Arbeiter-Zeitung, apolitical mass-paper Illustrierte Kronen Zeitung and German-nationalist Deutsches Volksblatt (Paupié, 1960; Melischek & Seethaler, 2016). Sentences were automatically annotated with topics and classified into four analytical categories: MIG, MIN, FAS and CTX, reflecting scholarship on migration, minorities and assimilation in Austria (Brix, 2001; Komlosy, 2004; Hahn, 2000; Steidl, 2020).
Seven German transformer models were fine-tuned for sentiment analysis using the SentiAnno 2.0 gold standard corpus (Brozić, 2025b). These models included contemporary ones, such as GottBERT (Scheible et al., 2024) and bert-base-german-cased (MDZ Digital Library Team, 2019), as well as historical ones, including hmBERT (Schweter et al., 2022) and GHisBERT (Beck & Köllner, 2023), and BERT and ELECTRA models pre-trained on Europeana historical newspapers (Schweter, 2020).
The original SentiAnno annotation scheme included four sentiment labels: “negative” , “neutral”, “positive” and “mixed”. However, initial experiments showed that including the rarely occurring, conceptually ambiguous “mixed” label substantially reduced model accuracy because the models failed to learn consistent patterns from these cases. Therefore, with the help of annotator comments, the “mixed” annotations were consolidated into the sentiment category they leaned towards, improving model accuracy.
The resulting corpus comprises 2,594 annotated text segments labeled as “negative”, “neutral”, or “positive”, displaying a significant class imbalance: 49.8% are negative, 39.3% are neutral, and 10.9% are positive. The remaining class imbalance was mitigated via class-weighted loss and stratified splits. Despite being trained on contemporary German, GottBERT achieved the best performance (accuracy: 0.74; macro-F1: 0.68), outperforming historically pretrained models. This is likely because GottBERT was pretrained on substantially more data than the historical models (the German portion of the OSCAR dataset, 121 GB; Ortiz Suárez et al., 2019).
The confusion matrix (see Figure 1) of the fine-tuned SentiAnno-GottBERT model shows that negative sentiment is captured reliably, neutral reasonably well, while predicting positive sentiment remains difficult. Sarcasm, irony and ideologically affirmative yet lexically aggressive language (for instance in antisemitic discourse) remain challenging.
Figure 1. GottBERT confusion matrix
Applied to MigraAnno, the model confirms that discourse on migrants and minorities is predominantly negative and neutral (Hampton, 2008; de Veen & Thomas, 2022). Over time, sentiment trajectories correlate with political alignment: low polarity aligns with censorship or conformity, while high polarity marks ideological activism.
In Wienerisches Diarium, MIN topics on Jews, Turks and Croats show strong negative sentiment persisting after the 1781 Edict of Toleration, while FAS topics on multilingualism and German display positive sentiment aligned with Enlightenment ideals (Pesalj, 2019; Scott, 1990; Klingenstein, 1993; Evans, 2004). In the censored early 19th century, Österreichischer Beobachter and Der Wanderer reported negatively on migration, exiles and revolutionary movements (Emerson, 1968; Heyck, 1989). Late 19th-century newspapers diverged: Neue Freie Presse showed strong negative sentiment toward Slavic minorities while expressing negative sentiment toward antisemitism, whereas Catholic Das Vaterland showed negative sentiment toward Jews but more neutral to mildly positive sentiment toward Slavic Catholics. In the early 20th century, Arbeiter-Zeitung combined negative sentiment toward policies driving emigration with positive sentiment toward refugees, Deutsches Volksblatt exhibited persistently negative and exclusionary sentiment, and Illustrierte Kronen Zeitung remained largely neutral before turning increasingly positive toward National Socialism.
A close reading revealed a core limitation of sentence-level polarity. “Negative” sentiment sometimes reflected attitudes toward government policy while implicitly expressing empathy toward migrants or minority groups. For instance, the negativity expressed by Arbeiter-Zeitung regarding "Galicia" was directed at governmental shortcomings and demonstrated solidarity with refugees. In contrast, the negativity exhibited by Deutsches Volksblatt was directed at refugees. Similarly, the negativity displayed by Neue Freie Presse regarding the topic “Jews” often signified criticism of antisemitism. Aspect-based sentiment analysis and stance detection may offer more fine-grained alternatives in the future (Dejaeghere et al., 2024; Hamdi et al., 2021). Improvements to model accuracy could include domain-adaptive pretraining on historical corpora, expanding positive and figurative examples, and reviewing borderline cases with experts, as in active learning (Brunner et al., 2020; Schmidt et al., 2021; Prabhu et al., 2021).
This presentation contributes to the growing body of work on sentiment analysis in the Digital Humanities and presents novel fine-tuned models for sentiment analysis of historical texts in German. The best-performing model, SentiAnno-GottBERT (Brozić, 2025c), is openly available on HuggingFace, while the corpora are available on Zenodo, following FAIR principles (Wilkinson et al., 2016). Together, these findings underscore the analytical potential and limits of sentiment analysis in historical newspaper research, highlighting the necessity of integrating computational methods with contextual and source-critical interpretations.