DH 2026

Daejeon, July 27–31

Poster

MigraAnno: Creating a topic-specific corpus from a large newspaper collection

Lucija Brozić (Krušić)
University of Graz, Austria · lucija.krusic@uni-graz.at

This poster introduces MigraAnno, a topic-specific subcorpus of migration- and  minority-related discourses from Austrian historical newspapers published between 1703 and 1938. The base corpus includes the monarchistic Wienerisches Diarium and conservative Österreichischer Beobachter, liberal Der Wanderer and Neue Freie Presse, the Catholic-conservative Das Vaterland, the social democratic Arbeiter-Zeitung, the apolitical Illustrierte Kronen Zeitung and the German-nationalist, anti-semitic Deutsches Volksblatt. They were scraped from the DIGITARIUM (Resch et al., 2021) and ANNO (Österreichische Nationalbibliothek, 2021) repositories and post-corrected using the fine-tuned hmbyt-NewsEyeAnno OCR model (Krušić, 2024).

Creating a specialized, topic-specific corpus is essential for answering research questions in the Digital Humanities using qualitative and quantitative analyses. This project addresses the challenge of identifying thematically relevant material in historical newspaper collections, which is difficult due to OCR errors, implicit discourse, inconsistent spelling, and a lack of descriptive metadata (Pfanzelter et al., 2021; Oberbichler & Pfanzelter, 2022). To overcome these obstacles, MigraAnno was created by integrating seeded BERTopic modeling (Grootendorst, 2022), qualitative topic evaluation, and transformer-based relevancy classification. The goal was to develop a reliable, contextually grounded corpus suitable for further historical, linguistic, and computational analysis.

Since historical newspapers frequently incorporate discussions of migration and minority discourse within broader topics such as labor, language, religion, education, and wars (Krušić Brozić, 2025; King & Wood, 2000; Hahn, 2000), context-aware modeling was crucial for capturing implicit discourse. BERTopic was chosen because it uses transformer embeddings and can incorporate seed words, or domain-relevant keywords. This allows relevant topics to be identified while maintaining the exploratory nature of unsupervised learning (Grootendorst, 2024). A separate BERTopic model was trained for each of the eight newspapers across four historical periods. Sentences were used as the modeling unit to maximize granularity. Seed words were organized into categories representing migration terms, minority groups, factors influencing assimilation (e.g., education, religion, and labor), and historical contexts and catalysts of migration (e.g., wars, movements, and epidemics). For each newspaper, the resulting topics were examined through a human-centered evaluation process. The selection of relevant topics was based on semantic coherence, alignment between representative documents and topic labels, and demonstrable relevance to migration- and minority-related discourse.

BERTopic produced 206,557 candidate sentences across all newspapers. However, a qualitative inspection revealed that some sentences were clearly related to migration or minority discourse, while others, such as administrative notices or travel price announcements, shared vocabulary with relevant topics but were irrelevant to the research objectives. These findings demonstrated that topic membership alone was insufficient for defining the corpus, which motivated the integration of a supervised relevancy classification step.

To this end, a stratified sample was manually annotated as either relevant or irrelevant. This sample was then used to fine-tune the GottBERT transformer-based model (Scheible et al., 2024) for relevancy classification, resulting in the fine-tuned MigraAnno-GottBERT-relevancy model (Brozić, 2025a). The model achieved a macro F1 score of 0.93, demonstrating its ability to capture contextual cues of historical German. The model was then applied to all candidate sentences to filter out those that were contextually irrelevant. Consequently, the final MigraAnno corpus comprises 96,871 relevant sentences enriched with metadata, including newspaper title, date, topic label, contextual sentences, category, and relevancy probability.

To mitigate bias and ensure representativeness, it is essential to be transparent about the data, tools, and contextualization of sources, which includes providing metadata and information about the datafication process (Oberbichler & Pfanzelter, 2022). All corpus files, topic models, and the relevancy classifier are openly available via GitHub (Brozić, 2025c; Brozić 2025e), Zenodo (Brozić 2025b), and HuggingFace (Brozić 2025a; Brozić 2025d), promoting FAIR principles (Wilkinson et al., 2016). MigraAnno shows that a focused, interpretable, and reusable, topic-specific corpus can be created from a large, noisy historical newspaper collection using a pipeline that combines transformer-based topic modeling, supervised classification, and human evaluation.

The MigraAnno corpus supports a wide range of Humanities and Digital Humanities analyses, enabling close reading and specialised investigations of migration, minorities, and related social, political, and cultural contexts. It also provides a robust training resource for fine-tuning classifiers and developing semi-supervised or supervised topic models for historical German texts.

References
  1. Brozić, Lucija (2024): hmbyt-NewsEyeAnno. https://huggingface.co/lukru/hmbyt-NewsEyeAnno.
  2. ---. 2025a. Topic models. https://huggingface.co/lukru/models (zugegriffen: 24. November 2025).
  3. ---. 2025b. lukru/MigraAnno-GottBERT-relevancy · Hugging Face. huggingFace. https://huggingface.co/lukru/MigraAnno-GottBERT-relevancy (zugegriffen: 24. November 2025).
  4. ---. 2025c. MigraAnno-relevancy-classification. https://github.com/lucijakrusic/MigraAnno-relevancy-classificationhttps://github.com/lucijakrusic/MigraAnno-relevancy-classification.
  5. ---. 2025d. MigraAnno. Python. 20. November. https://github.com/lucijakrusic/MigraAnno.
  6. ---. 2025e. MigraAnno. Zenodo, 20. November. doi:10.5281/ZENODO.17659577, https://zenodo.org/doi/10.5281/zenodo.17659577 (zugegriffen: 20. November 2025).
  7. Grootendorst, Maarten. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv, 11. März. http://arxiv.org/abs/2203.05794 (zugegriffen: 28. Februar 2024).
  8. Grootendorst, Marteen (2024): Seed Words - BERTopic. Documentation. BERTopic getting started. https://maartengr.github.io/BERTopic/getting_started/seed_words/seed_words.html (zugegriffen: 22. September 2025).
  9. Hahn, Sylvia(2000): Inclusion and Exclusion of Migrants in the Multicultural Realm of the Habsburg „State of Many Peoples“. Histoire sociale / Social History (1. November). https://hssh.journals.yorku.ca/index.php/hssh/article/view/4568 (zugegriffen: 11. Februar 2024).
  10. King, Russell und Nancy Wood, Hrsg. (2001): Media and migration: constructions of mobility and difference. Routledge research in cultural and media studies 8. London ; New York: Routledge.
  11. Krušić Brozić, Lucija. 2025. Discovering a Shared Past Topic Modelling of Austrian Historical Newspapers. In: Monografija u povodu 20 godina rada Odjela za informacijske znanosti u Zadru, hrsg. von Mirko Duić, Mate Juric, und Nikolina Peša Pavlović. University of Zadar. doi:10.15291/9789533315638, https://morepress.unizd.hr/books/index.php/press/catalog/book/145 (zugegriffen: 20. November 2025).
  12. Krušić, Lucija (2024): hmbyt-NewsEyeAnno. https://huggingface.co/lukru/hmbyt-NewsEyeAnno.
  13. Oberbichler, Sarah, Emanuela Boroş, Antoine Doucet, Jani Marjanen, Eva Pfanzelter, Juha Rautiainen, Hannu Toivonen und Mikko Tolonen. 2022. Integrated interdisciplinary workflows for research on historical newspapers: Perspectives from humanities scholars, computer scientists, and librarians. Journal of the Association for Information Science and Technology 73, Nr. 2 (Februar): 225–239. doi:10.1002/asi.24565, https://asistdl.onlinelibrary.wiley.com/doi/10.1002/asi.24565 (zugegriffen: 22. Juli 2025).
  14. Oberbichler, Sarah und Eva Pfanzelter (2021): Topic-specific corpus building: A step towards a representative newspaper corpus on the topic of return migration using text mining methods. Journal of Digital History, Nr. jdh001. https://journalofdigitalhistory.org/en/article/4yxHGiqXYRbX.
  15. Scheible, Raphael, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer und Martin Boeker (2024): GottBERT: a pure German Language Model. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, hrsg. von Yaser Al-Onaizan, Mohit Bansal, und Yun-Nung Chen, 21237–21250. Miami, Florida, USA: Association for Computational Linguistics, November. doi:10.18653/v1/2024.emnlp-main.1183, https://aclanthology.org/2024.emnlp-main.1183/ (zugegriffen: 3. September 2025).
  16. Wilson Black, Joshua (2023): Creating specialized corpora from digitized historical newspaper archives. Digital Scholarship in the Humanities 38, Nr. 2 (31. Mai): 779–797. doi:10.1093/llc/fqac079, https://academic.oup.com/dsh/article/38/2/779/6957053 (zugegriffen: 14. Juni 2023).