DH 2026

Daejeon, July 27–31

Poster

A Preliminary Study on the Cultural Bias in Indexing Practices in a Large 19th Century Colonial Archive

Sebastiaan Peeters
University of Twente, Netherlands, The · sebastiaan.peeters@utwente.nl
Annemieke Romein
University of Twente, Netherlands, The; Walter-Benjamin Kolleg, University of Bern, Bern, Switzerland; READ-COOP SCE, Innsbruck, Austria · c.a.romein@utwente.nl
Andreas Weber
University of Twente, Netherlands, The · a.weber@utwente.nl

Recent advances in Automated Text Recognition (ATR) (Pinche and Stokes 2024) have encouraged cultural heritage institutions to initiate mass digitization projects, which has led to a deluge of new transcription data (Thylstrup 2024). However, these transcribed images, with transcription errors and challenging historical handwriting, often end up collecting dust on their virtual shelves due to a lack of accessibility (Nockels et al. 2024). The traditional way of unlocking such archives is accessing them through their historical organization. This often entails navigating some form of index or glossary which offers an access point to the remainder of the archives. Such indices and glossaries are very dense in information and offer insights in the contents of the full archives, as pointed out by Colavizza et al. (2019). They developed a framework for digitizing indices based on their experience with the Venice city archives. However, it should be noted that such historical forms of access and organization were also shaped by the preoccupations and cultural context of their creators. Therefore, they reflect a strong cultural bias, and as such may omit information that modern users of the archive might find pertinent (Ortolja-Baird and Nyhan 2022).

This poster will explore the contents and cultural biases of the glossaries of the archives of the Dutch Ministry of Colonies from 1850-1900 (NL-HaNa 2.10.02). The archive contains records of the daily affairs of the ministry, which, after the dissolution of the Dutch West (1792) and East (1800) Indies Companies, assumed control over all Dutch colonies. The Ministry of Colonies archives was recently digitized and all ca. five million scans were transcribed using Loghi (van Koert et al. 2024). Like in other Dutch ministerial archives of its kind, archive navigation starts from glossaries called klappers, which link to indices, which in turn link to the main dossiers, verbalen (Otten 2004). Therefore, the glossary contents determine what is findable in the traditional search workflow (see figure 1).

Figure Search workflow of the Ministry of Colonies Archive

As the Dutch Ministry of Colonies archive is entirely constructed by Dutch government officials, it was intended to serve the needs of the colonial administration. Consequently, their biases are reflected in the organization of the search workflow. As Jeurgens and Karabinos (2020) point out, such record keeping systems are in themselves a continuation of colonial power dynamics by controlling access to information. Ann Laura Stoler demonstrates this for this particular archive through several case studies in her book Along the Archival Grain, which demonstrate how the insecurities of a colonial ruling class shaped their decision making in which information to include in the archives (Stoler 2009).

Our study of the glossaries of the archives of the Dutch Ministry of Colonies and their biases therefore proceeds in two steps. First we will offer an overview of what is contained in the glossaries in terms of content type (i.e. LLM based named entity categorization: person, ship, place, product, … ) and if there are any evolutions in indexing practices over time, for example as a result of the increased document production towards the end of the 19th century (Vriend 2025). This inquiry is part of and builds on a larger and still ongoing research project in which we are developing a navigational tool for unlocking the contents of the archive by digitally recreating the historical search workflow (Peeters et al. 2025). As part of this larger project, the glossaries, which in the historical navigation workflow form the primary access points to the archives, have been extracted and segmented into individual entries.

In a second step of our analysis of cultural bias we will, using select case studies from within the archive, compare personal names contained in the full archive text to those that made it into the glossaries, exploring patterns in the omissions. In this way, it also applies a form of tool criticism, which critically considers the implications of following the historical search workflow for navigating archives.

While the exploration of the glossaries in isolation will cover the full breadth of the archive, the comparison between their contents and those of the archive proper is only a preliminary case study on a selection within the archive that will rely on manual linking and either manual or semi-automated tagging and disambiguation. Once the navigation tool referenced above is complete, a large-scale study covering the entire archive will be conducted using fully automated NER tagging. On top of presenting the contents of the Ministry of Colonies archive, the poster demonstrates the value of research into historical finding aids and offers some preliminary results which may encourage additional research in this direction.

References
  1. Colavizza, Giovanni / Ehrmann, Maud / Bortoluzzi, Fabio (2019): “Index-Driven Digitization and Indexation of Historical Archives.” Frontiers in Digital Humanities 6. DOI: 10.3389/fdigh.2019.00004.
  2. Jeurgens, Charles / Karabinos, Michael (2020): “Paradoxes of Curating Colonial Memory.” Archival Science 20(3), 199–220. DOI: 10.1007/s10502-020-09334-z.
  3. Koert, Rutger van / Klut, Stefan / Koornstra, Tim / Maas, Martijn / Peters, Luke (2024): “Loghi: An End-to-End Framework for Making Historical Documents Machine-Readable.” In: Mouchère, Harold / Zhu, Anna (eds.): Document Analysis and Recognition – ICDAR 2024 Workshops. Cham: Springer Nature Switzerland. DOI: 10.1007/978-3-031-70645-5_6.
  4. Nockels, Joseph / Gooding, Paul / Terras, Melissa (2024): “The Implications of Handwritten Text Recognition for Accessing the Past at Scale.” Journal of Documentation 80(7), 148–167. DOI: 10.1108/JD-09-2023-0183.
  5. Ortolja-Baird, Alexandra / Nyhan, Julianne (2022): “Encoding the Haunting of an Object Catalogue: On the Potential of Digital Technologies to Perpetuate or Subvert the Silence and Bias of the Early-Modern Archive.” Digital Scholarship in the Humanities 37(3), 844–867. DOI: 10.1093/llc/fqab065.
  6. Otten, F. J. M. (2004): Gids voor de archieven van de ministeries en de hoge colleges van staat, 1813–1940. Available at: https://resources.huygens.knaw.nl/retroboeken/archiefgids_overheid/#page=0&accessor=toc_1&view=homePane.
  7. Peeters, Sebastiaan / Romein, C. Annemieke / Weber, Andreas (2025): “Towards a Navigator Tool for Dutch Verbaal-Archives: Leveraging Nineteenth-Century Archival Logic for Keyword Search.” In: Arnold, Taylor / Fantoli, Margherita / Ros, Ruben (eds.): Computational Humanities Research 2025. Anthology of Computers and the Humanities. DOI: 10.63744/ByKY4LADHCBb.
  8. Pinche, Ariane / Stokes, Peter (2024): “Historical Documents and Automatic Text Recognition: Introduction.” Journal of Data Mining & Digital Humanities, March, Article 13247. DOI: 10.46298/jdmdh.13247.
  9. Stoler, Ann Laura (2009): Along the Archival Grain: Epistemic Anxieties and Colonial Common Sense. Princeton: Princeton University Press.
  10. Thylstrup, Nanna Bonde (2024): The Politics of Mass Digitization. Cambridge, MA: MIT Press.
  11. Vriend, Nico (2025): “An Archive in Numbers: The Pulse of the Dutch Ministry of Colonies, 1813–1900.” Archival Science 25(2), Article 17. DOI: 10.1007/s10502-025-09483-z.