Daejeon, July 27–31
The boundary between administrative and literary writings is seldom more blurred than in Jans teesteye. This literary work by the influential Middle Dutch author Jan van Boendale (c. 1280-1351?) opens as follows (vs. 1-8):
| Alle die ghene die dit werc | To everyone who |
| Sien lesen ende horen | sees, reads, orhears this work |
| Die gruetic Jan gheheten Clerc | I greet you, Jan, called the clerk, |
| Vander Vueren gheboren | born in Tervuren |
| Boendale heetmen mi daer. | where they call me Boendale. |
| Ende wone te Andwerpen nu | I now live in Antwerp |
| Daer ic ghescreven hebbe menech jaer | where I have written for several years |
| Der scepenen brieve dat segghic u | the deeds, that is what I tell you. |
Boendale states that he has written the Antwerp deeds – aldermen’s letters often recording property transactions – for several years. He reaffirms this in the opening verses of the work, which echo a typical opening formula of Middle Dutch deeds (vs. 1-2). His position as a clerk is further confirmed in the Antwerp municipal accounts (Gordeau 2020). Many Antwerp deeds still survive today, yet, surprisingly, no attempt has been made to identify which deeds were written down by Jan van Boendale himself.
This is remarkable given that Boendale was also an exceptionally prolific literary author (Reynaert, 2002; Van Oostrom, 2013). He produced a substantial oeuvre targeted at a lay audience, including patricians and aristocrats, thereby contributing significantly to the development of an emerging lay culture in the Dutch vernacular (Appelmans, 2002). Although the precise delineation of his oeuvre remains debated, it is consistently associated with the so-called ‘Antwerp School’, a group of eleven interrelated works composed in early fourteenth-century Antwerp. These texts are moralistic, didactic and historiographic, they reflect similar norms and values, and they contain recurring phrases, creating a tight intertextual network (e.g. De Vries, 1844-1848; Reynaert, 2002; Snellaert, 1869; Vandyck, 2025). Boendale certainly authored at least part of the Antwerp School texts, and recent scholarship favours the maximalist interpretation that he authored the entire corpus (Gordeau, 2020; Vandyck, 2025).
Developing a dataset of digitized and transcribed Antwerp deeds affords us a unique opportunity to connect a literary author to his day job, and to gain a rare insight into his daily life. This abstract presents a dataset of deeds that we collected for this purpose, and the research that is currently in progress.
The collected dataset comprises entries on 514 deeds written in Antwerp in Middle Dutch between c. 1300 and 1355, the period when Boendale was active as a municipal clerk. Compiling an exhaustive catalogue of the deeds is challenging, as many relevant archival holdings are incompletely or incorrectly catalogued but our catalogue nevertheless aims to be as comprehensive as possible. The dataset includes three metadata fields for each deed – date, repository, and a unique identifier – and each document will be linked to a digital reproduction as well as two transcriptions (Figure 1).
Our dataset builds on a handlist compiled by former archivist Jos Van den Nieuwenhuizen. Archivist Michel Oosterbosch further enriched this handlist with metadata and a few more entries. Of the 453 deeds relevant to us, cross-referencing with a handlist by Godfried Croenen yielded 55 extra entries. We retrieved six more deeds during the process and added those too, resulting in a collection of 514 documents in total. For 470 we have a photographic reproduction and these are thus the most relevant – it was not feasible to digitise the remainder of the deeds within the scope of the project. There are five different sources of facsimiles for the documents. The core is formed by the reproductions assembled by Van den Nieuwenhuizen, which were scanned at the Antwerp State Archives in 2020 under the direction of Oosterbosch. Other reproductions were made by Chris Dewulf, Kamiel Temmerman, FelixArchief and Stadsarchief Mechelen. Some deeds are available in multiple digital facsimiles, which we linked in the dataset by assigning a unique identifier to each deed. We selected the best image (“Final”) for each deed.
We transcribed a subset of deeds manually and diplomatically following the guidelines in Haverals and Kestemont (2023) (Figure 2). We trained a Handwritten Text Recognition (HTR) model based on our transcriptions on Transkribus, obtaining a CER of 2.62% on the control set. The model will be released together with the final dataset. To expand the abbreviations in the diplomatic transcriptions, we are currently in the process of finetuning a transducer, previously proposed for this task by Haverals and Kestemont (2023). We will share the full dataset – including scans, metadata, transcription layers, and models – on Zenodo upon finalization. The dataset can already be explored in its current form via a static page, which was made with Necturus (Shahidzadeh Asadi, 2025). https://mikekestemont.github.io/deeds-of-antwerp/#/collection/1
Figure 1: Excerpt of the metadata, with columns showing the unique identifier, date, repository and shelfmark, and the different transcriptions and images. “Final” refers to the best image, renamed after the ID of the deed.
Figure 2: Deed kept in Stadsarchief Mechelen (RAA, Pitzemburg) dated 17-06-1338 alongside its manual diplomatic transcription. Based on the dating of this hand, it could be Jan van Boendale’s.
An extra layer, soon to be added to the metadata, is the identification of the scribal hands in the dataset. Since the deeds are unsigned, paleographic and chronological features are the primary means of distinguishing different hands. Our approach to this identification was a combination of automatic and human annotation of the data.
The state of the art in computational handwriting identification was developed by Raven et al. (2024), who helped us set up the project. They rely on unsupervised pretraining for feature learning. We applied their pretraining routines with minor modifications using a Vision Transformer (small-ViT), with AttMask for self-supervised training. After pretraining, we performed a preliminary clustering of the deeds, by extracting fixed-length features vector for the individual documents (using VLAD-encoding), reducing dimensionality via PCA and UMAP and performing clustering with HDBSCAN (Figure 3). Manual inspection of the resulting clusters already suggested convincing groupings, characterized by high script similarity. Interestingly, the local structure of the embedding space was more convincing than the global structure: deeds assigned to the same cluster were invariably very close in script, but some of the (e.g. adjacent) clusters might require merging.
Figure 3: Preliminary clustering of the deeds after pretraining, made by extracting fixed-length features vector for the individual documents (using VLAD-encoding), reducing dimensionality via PCA and UMAP and performing clustering with HDBSCAN.
To validate these clusters, we organized an experimental community annotation event in which volunteer medievalists (historians, linguists, paleographers, literary scholars, etc.) participated. Two main annotation tasks were designed. In the first, annotators were presented with triplets of documents and asked to identify the odd one out — the deed least similar in handwriting to the other two. The triplets were based on the aforementioned clusters. In the second, each annotator was assigned one of the clusters identified by the preliminary automatic analysis, from which they had to “purge” the outliers. In the next step, the clusters were collectively evaluated for potential merging. The results of both tasks will be used to fine-tune the model in a later stage. In this talk, we will present the comparison between the human and automated clustering.
Caroline Vandyck’s research is funded by Research Foundation Flanders (FWO) and is a result of her Fellowship fundamental research “Boendale’s Many Faces: Modeling Historical Positionality in Jan van Boendale’s (c. 1280-1351?) Oeuvre” (1133925N). Godfried Croenen’s research is funded by the FWO Senior research project “Neverending Stories: Unlocking Complicated Textual Traditions and Historiographical Networks. Showcase: The Production and Dissemination of Boendale's Chronicle of the Dukes of Brabant ('Brabantsche yeesten')” (G025725N).