DH 2026

Daejeon, July 27–31

Poster

Behind the Labels: Annotating 100 Years of Historical Job Advertisements for NER

Klara Venglarova
University of Graz, Austria · klara.venglarova@uni-graz.at
Raven Adam
University of Graz, Austria · raven.adam@uni-graz.at
Wiltrud Mölzer
University of Graz, Austria · wiltrud.moelzer@uni-graz.at
Jörn Kleinert
University of Graz, Austria · joern.kleinert@uni-graz.at
Georg Vogeler
University of Graz, Austria · georg.vogeler@uni-graz.at

Job advertisements from historical newspapers offer a rich resource for studying labor markets, but extracting structured information from them remains challenging due to their unstructured format, OCR errors, complex layouts, and archaic language—challenges that are common to Named Entity Recognition on historical texts more generally (Novotny et al., 2023; Avram et al., 2024; Schweter et al., 2022). Such challenges, as well as the research potential of these sources, have also been identified by several scholars working with job advertisements or digitized newspapers more broadly (Bunout, 2023; Håkansson et al., 2023; Oberbichler & Pfanzelter, 2023; Schulz et al., 2014; Wevers, 2022; Wevers & Smits, 2020).

Within the JobAds project, we developed a comprehensive pipeline for processing Austrian newspapers (Österreichische Nationalbibliothek, 2021) from 1850 to 1950, focusing on job advertisements. The pipeline encompasses newspaper pages segmentation, OCR, text-type classification, and training of machine learning models for NER. The NER component relies on an ELECTRA-based architecture fine-tuned on historical German newspaper data ('zeitungs-lm-v1' (Schweter, 2024)) to better capture diachronic linguistic variation typical of the period. Details of the pipeline architecture and model training are described in Adam et al. (2025) and Venglarova et al. (2024).

While these prior publications have presented our results in extracting job titles and occupational data, this poster shifts the focus to a crucial yet often underrepresented aspect of the pipeline: the manual annotation process used to train a Named Entity Recognition (NER) model to identify job titles, requirements, skills, salary information, and other job-relevant entities. The structured data produced by the pipeline enables analyses of long-term developments in job demand and supply; for a detailed study see Mölzer and Kleinert (2024).

To generate high-quality training data for the NER model, the project team, with the support of student assistants, manually annotated a representative subset of 14,985 job advertisements, using the doccano open-source annotation tool (Nakayama et al., 2018). The subset of the JobAds corpus chosen for annotation was derived through disproportionate stratified sampling with respect to newspaper title and year. The aim was twofold: to provide a small, high-quality annotated dataset and to create the training material necessary for large-scale automated extraction of job-relevant information.

Annotations covered several layers of information, starting from layout when we were annotating visual blocks on the newspaper page. Further annotations, already on the text level, included individual components of job advertisements, such as job titles, salary information, and requirements.

The process revealed the inherent challenges of translating human expression into structured categories. For instance, in the case of visual annotations of job advertisements and their classification into sub-categories (job offers, job searches, service offers, mediations and headings), syntax was sometimes ambiguous: a phrase like “Perfekte Kontoristin sucht Schuhfabrik” [ Perfect office clerk seeks shoe factory or Shoe factory seeks a perfect office clerk ] leaves it unclear whether it is a job offer or a job search, complicating classification, and the ‘correct solution’ is only available based on the context – it is more probable that a shoe factory is looking for an office clerk than vice versa. Similarly, sometimes it is unclear, whether a segment should be classified as a service offer, or if it is not a job advertisement at all and just advertising one’s business, such as in the short text “Massage” followed by an address.

Another problem of this type was a missing historical context, such as in the advertisement “Madelaine Arztesgattin” [doctor’s wife] followed by the address, as it is not clear without additional research, whether Madelaine is a name, a historical name of a position, or simply a product.

Similar problems appear when annotating entities within the text. Not all phrases clearly fit predefined labels: job titles could be literal descriptions of persons, such as “Junger Mann” [young man] , “Französinnen” [French women] , or “Mädchen” [maid] , requiring the addition of a new label for these “not-positions.”

Through iterative refinement of annotation guidelines and continuous discussion among annotators, the team developed strategies to handle such ambiguities, increasing consistency and transparency. While many cases required interpretation based on the specific historical and textual context of each advertisement, recurring patterns were incorporated into the evolving annotation guidelines. This allowed the team to balance context-sensitive judgment with consistent rule-based decisions across the dataset. These examples illustrate how seemingly small annotation decisions can have significant downstream effects on model behaviour.

The resulting dataset not only enabled the training of reliable NER models but also provided methodological insights into the complexities of annotating historical texts, emphasizing the human judgment that underpins automated analysis. Based on these rigorous and methodological guidelines the resulting training data allowed the creation of an NER model that achieved a F1 score of 0.95 on a held out test dataset with a human-corrected text, and a F1 score of 0.83 on a completely unrelated dataset of German newspapers with a messy OCR (Adam et al., 2025).

References
  1. Adam, R., Venglarova, K., & Vogeler, G. (2025). Exploring Historical Labor Markets: Computational Approaches to Job Title Extraction. Journal of Data Mining & Digital Humanities, NLP4DH. https://doi.org/10.46298/jdmdh.15038
  2. Avram, A.-M., Iuga, A., Manolache, G.-V., Matei, V.-C., Micliuş, R.-G., Muntean, V.-A., Sorlescu, M.-P., Şerban, D.-A., Urse, A.-D., Păiş, V., & others. (2024). Histnero: Historical named entity recognition for the romanian language. International Conference on Document Analysis and Recognition , 126–144. 
  3. Bunout, E. (2023). Contextualising queries: Guidance for research using current collections of digitised newspapers. Digitised Newspapers: A New Eldorado for Historians, 277–300.
  4. Håkansson, P. G., Karlsson, T., & La Mela, M. (2023). Running out of time: Using job ads to analyse the demand for messengers in the twentieth century. Scandinavian Economic History Review, 71(3), 299–318.
  5. Mölzer, W., & Kleinert, J. (2024). Emergence of the Austrian labor market. https://static.uni-graz.at/fileadmin/_files/_project_sites/_historical-job-ads/Emergence_Austrian_labor_market.pdf
  6. Nakayama, H., Kubo, T., Kamura, J., Taniguchi, Y., & Liang, X. (2018). doccano: Text Annotation Tool for Human. https://github.com/doccano/doccano
  7. Novotny, V., Luger, K., Štefánik, M., Vrabcova, T., & Horak, A. (2023). People and Places of Historical Europe: Bootstrapping Annotation Pipeline and a New Corpus of Named Entities in Late Medieval Texts. In A. Rogers, J. Boyd-Graber, & N. Okazaki (Eds), Findings of the Association for Computational Linguistics: ACL 2023 (pp. 14104–14113). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-acl.887 
  8. Oberbichler, S., & Pfanzelter, E. (2023). Tracing Discourses in Digital Newspaper Collections. In E. Bunout, M. Ehrmann, & F. Clavert (Eds), Digitised Newspapers – A New Eldorado for Historians? (pp. 125–152). https://doi.org/10.1515/9783110729214-007
  9. Österreichische Nationalbibliothek. (2021). ANNO Historische Zeitungen und Zeitschriften. https://anno.onb.ac.at/
  10. Schulz, W., Maas, I., & Van Leeuwen, M. H. (2014). Employer’s choice–Selection through job advertisements in the nineteenth and twentieth centuries. Research in Social Stratification and Mobility, 36, 49–68.
  11. Schweter, S. (2024). Zeitungs-lm-v1 (Revision cf6726c) . Hugging Face. https://doi.org/10.57967/hf/3176 
  12. Schweter, S., März, L., Schmid, K., & Çano, E. (2022, September 5). hmBERT: Historical Multilingual Language Models for Named Entity Recognition
  13. Venglarova, K., Adam, R., & Vogeler, G. (2024). Extracting position titles from unstructured historical job advertisements. In M. Hämäläinen, E. Öhman, S. Miyagawa, K. Alnajjar, & Y. Bizzoni (Eds), Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities (pp. 75–84). Association for Computational Linguistics. https://aclanthology.org/2024.nlp4dh-1.8
  14. Wevers, M. (2022). Mining Historical Advertisements in Digitised Newspapers. In Digitised Newspapers – A New Eldorado for Historians? (pp. 227–252). De Gruyter Oldenbourg. https://doi.org/10.1515/9783110729214-011
  15. Wevers, M., & Smits, T. (2020). Detecting Faces, Visual Medium Types, and Gender in Historical Advertisements, 1950–1995 (pp. 77–91). https://doi.org/10.1007/978-3-030-66096-3_7