DH 2026

Daejeon, July 27–31

Thu, July 3015:20–16:20S090101-102
Short Paper

Outside the Box: Building a Pipeline for Extracting Demographic Data from Late Ottoman Census Records

Olaf Berg
Ruhr-Uni-Bochum, Institut f. Arabistik, ERC LOOP, Germany · olaf.berg@rub.de
Johann Büssow
Ruhr-Uni-Bochum, Institut f. Arabistik, ERC LOOP, Germany · Johann.Buessow@ruhr-uni-bochum.de
Sarah Büssow
Ruhr-Uni-Bochum, Institut f. Arabistik, ERC LOOP, Germany · sarah.buessow@rub.de
Liat Bonen
University of Haifa, Jewish Studies, eLijah-Lab for Digital Humanities · lbonen@staff.haifa.ac.il
Moshe Lavee
University of Haifa, Jewish Studies, eLijah-Lab for Digital Humanities · mlavee@univ.haifa.ac.il

This paper addresses the technical challenges of applying handwritten text recognition (HTR) to demographic data preserved in late Ottoman census registers. Two issues are central. First, the registers are handwritten in Ottoman Turkish, an under-resourced historical language for which effective pre-trained HTR models remain scarce. Second, the data consists of handwritten entries embedded in tabular forms, many with empty cells and others containing multi-line content that challenges standard layout assumptions (fig. 1). Existing table-segmentation methods perform poorly on such material, resulting in unreliable region detection and reduced transcription accuracy.

Figure Examples of Ottoman census records from Nüfus Book 21 (ISA Archive)
We speak from the perspective of end users of digital humanities tools, with a particular focus on the development, reuse, and sustainability of existing HTR software. Rather than proposing new algorithms, we combine established platforms and models and extend them through reusable Python scripts to construct a robust processing pipeline. Since few components functioned “out-of-the-box” for this material, the process required iterative experimentation and, at several points, deliberately “outside-the-box” technical decisions. We argue that documenting such user-driven adaptations is essential for improving the long-term sustainability and re-usability of DH infrastructures, particularly for under-resourced languages and complex historical documents.

The field

Since the 1990s, scholars have systematically extracted demographic and socio-economic data from surviving Ottoman census registers. Foundational studies include Duben and Behar’s analysis of Muslim and Armenian populations in Istanbul (1991), Cuno and Reimer’s work on the 1848 and 1868 Egyptian censuses (1997), and Alleaume and Fargues’s examination of nineteenth-century Egyptian census practices (1998). Later research broadened both scope and method: Saleh (2013) digitized late nineteenth-century Egyptian censuses, while Buessow (2011) reconstructed social characteristics of Palestinian regions. Sarah and Johann Buessow (2020) made largely invisible population groups visible through close reading of census registers. Buessow et.al. created a database of census materials from late Ottoman Gaza (2017-2023, 2024), and Campos (2018) analyzed social structures in Jerusalem using computed census data. The UrbanOccupationsOETR project extracted and published data on urban occupation from nineteenth-century Ottoman pre-census registers Kabadayı and Erünal (2024). All these datasets were produced through manual transcription, limiting scalability and reuse.

The present paper is part of the LOOP project (https://ercloop.hypotheses.org/loop), which aims to extract the complete corpus of late Ottoman census records of Palestine (467 volumes, 80,000+ pages, Berg and Büssow 2023) into a structured database enabling analyses from the micro to the macro level. These census records constitute a key part of the documentary heritage of the Eastern Mediterranean, underscoring the importance of sustainable digital access and processing pipelines.

In the past decade, HTR has advanced considerably, with two major platforms widely used in scholarly contexts: Transkribus, powered by the PyLaia engine, and the open-source platform eScriptorium, based on Kraken. Recent years have seen growing efforts to integrate HTR into Ottoman studies. Köse (n.d.) initiated the Ottoman Text Recognition Network in 2021 focusing on the creation of recognition models for the Transkribus platform. Kirmizialtın (2023) and the Digital Ottoman Corpora project published the first public generalized text recognition model for 19th-century Ottoman Turkish periodicals on Transkribus. The 2016 founded Open Islamicate Initiative develops HTR models for Arabic script on the eScriptorium platform. As base model for our pipeline, we used Smith et.al (2023) model created in the Automatic Collation for Diversifying Corpora project.

Typical HTR workflows consist of two separately trainable steps. First, document segmentation identifies regions, lines, and masks that represent the page layout. Second, text recognition converts the handwritten content within these masks into machine-readable text. Our census materials pose difficulties at both stages.

The segmentation challenge of tabular structures

The population registers from late Ottoman Palestine (c. 1880–1917) are handwritten within printed tabular forms. Many cells are empty, while others contain multiple lines of text, often skewed or overflowing cell boundaries. Repeated information, such as religious affiliation, is frequently indicated using ibid or ditto signs. Maintaining the tabular structure is therefore crucial for meaningful data extraction. Although Transkribus offers advanced table-detection features, these fail in our case due to the high proportion of empty cells. Existing algorithms typically infer the logical and physical table-structure by detecting text-blocks, an approach ill-suited to these registers.

The recognition challenge of an under-resourced language

HTR model availability and quality vary greatly by language. Ottoman Turkish, a predecessor of modern Turkish, is written in Arabic script and fell out of use following the language reform of 1932. The census registers employ Rika, a highly stylized variant of Arabic script. Recognition is further complicated by the multilingual nature of the data: many personal and place names derive from Arabic, Hebrew, Armenian, Greek, and other languages used in late Ottoman Palestine. At the start of our project in 2023, no public training sets existed for this combination of historical script and multilingual content. Our attempt to reproduce LLM-based table-recognition methods (Kim et. al. 2025) failed because current models lack proficiency in Ottoman Turkish.

Our “outside-the-box” pipeline

To address these challenges, we rethought segmentation entirely. Instead of attempting table recognition, we trained the segmentation model to draw a single line spanning the full width of each table row. During text recognition, vertical column boundaries are allowed to be recognized as pipe characters. In post-processing, the table is reconstructed from the text lines with the pipe character as a separator, as if from a CSV file. This approach achieved an overall text recognition rate of 94%, with an error rate of under 2% for years of birth. A further challenge emerged in segmentation: while line placement can be trained, the associated masks that define the text recognition area cannot. Due to empty cells and multi-line entries, these masks are often positioned too low. To address this, we export the segmented files and apply a custom Python script that generates new masks aligned to the full height of each table row. The corrected files are then re-imported for text recognition.

By foregrounding reuse, modular adaptation, and user-driven intervention within existing HTR infrastructures, this paper contributes a sustainable approach to processing under-resourced, trans-lingual documentary heritage, with relevance for similar materials preserved in large-scale archival collections.

References
  1. Alleaume, Ghislaine / Fargues, Philippe (1998): „La Naissance d’une Statistique d’État. Le Recensement de 1848 en Égypte“. Histoire & Mesure 13, 1/2: 147-193.
  2. Ben-Bassat, Yuval / Buessow, Johann (2024): Late Ottoman Gaza: An Eastern Mediterranean Hub in Transition. Cambridge: Cambridge University Press.
  3. Berg, Olaf, / Büssow, Johann (2023): “The ISA Corpus of Ottoman Census Registers, Or: What Sorts of Demographic Information Did the Ottoman Empire Collect in Palestine?”, in Research blog LOOP – Late Ottoman Palestinians, October 18. DOI: 10.58079/vxa9.
  4. Büssow, Johann (2011): Hamidian Palestine: Politics and Society in the District of Jerusalem 1872-1908. The Ottoman Empire and Its Heritage 46. Leiden, Boston: Brill.
  5. Buessow, Johann / Buessow, Sarah / Ben-Bassat, Yuval (2017-2023): Gaza Historical Database. <https://gaza.ub.rub.de/gaza/?p=home> [05. 05. 2026].
  6. Buessow, Sarah, / Johann Buessow (2020): “Domestic Workers and Slaves in Late Ottoman Palestine at the Moment of the Abolition of Slavery: Considerations on Semantics and Agency.”, in Slaves and Slave Agency in the Ottoman Empire, edited by Stephan Conermann / Gül Şen: 373-433. Bonn: V&R Unipress/Bonn University Press.
  7. Campos, Michelle U. (2018): “Placing Jerusalemites in the History of Jerusalem: The Ottoman Census (sicil-i nüfūs) as a Historical Source.” In Ordinary Jerusalem, 1840-1940. Leiden, Boston: Brill.
  8. Cuno, Kenneth M. / Michael J. Reimer (1997): “The Census Registers of Nineteenth-Century Egypt: A New Source for Social Historians.” British Journal of Middle Eastern Studies 24 (2): 193-216.
  9. Duben, Alan / Cem Behar (1991): Istanbul Households: Marriage, Family, and Fertility, 1880-1940. Cambridge studies in population, economy, and society in past time 15. Cambridge: Cambridge University Press.
  10. Kabadayı, M. Erdem / Efe Erünal (2024): “A Nineteenth-Century Urban Ottoman Population Micro Dataset: Data Extraction and Relational Database Curation from the 1840s Pre-Census Bursa Population Registers.”, in Scientific Data 11 (1): 570. DOI: 10.1038/s41597-024-03381-2.
  11. Kim, Seorin / Baudru, Julien / Ryckbosch, Wouter / Bersini, Hugues / Ginis. Vincent (2025): “Early Evidence of How LLMs Outperform Traditional Systems on OCR/HTR Tasks for Historical Records.”, preprint, arXiv, January 20. DOI: 10.48550/arXiv.2501.11623.
  12. Kirmizialtin, Suphan (2023): “From Script to Digital – Transforming Ottoman Turkish Texts with Suphan Kirmizialtin.” Transkribus Blog, November 13. <https://blog.transkribus.org/en/transforming-ottoman-turkish-texts-with-suphan-kirmizialtin> [05. 05. 2026].
  13. Köse, Yavuz (n.d.): “Ottoman Text Recognition Network.” Institutional page, <https://otrn.univie.ac.at/> [05. 05. 2026].
  14. OpenITI (n.d.): “About.” Open Islamicate Texts Initiative, <https://openiti.org/about.html> [05. 05. 2026].
  15. Saleh, Mohamed (2013): “A Pre-Colonial Population brought to Light: Digitization of the Nineteenth Century Egyptian Censuses.” Historical Methods 46 (1): 5-18.
  16. Smith, David A. / Murel, Jacob / Parkes Allen, Jonathan / Miller, Matthew T. (2023): “Automatic Collation for Diversifying Corpora: Commonly Copied Texts as Distant Supervision for Handwritten Text Recognition.” In Proceedings of the Computational Humanities Research Conference 2023, edited by Artjoms Šeļa, Fotis Jannidis, and Iza Romanowska, vol. 3558. CEUR Workshop Proceedings. CEUR. <https://ceur-ws.org/Vol-3558/paper1708.pdf> [05. 05. 2026].