DH 2026

Daejeon, July 27–31

Thu, July 3015:20–16:20S004104
Long Paper

Co-Designing an Annotation Infrastructure for Small, Interpretive Datasets with Arkindex

Alice Truc
Université de Montréal, Canada; Université Rennes 2, France · alice.truc@umontreal.ca
Léa Maronet
Centre National de la Recherche Scientifique (CNRS), France; École Pratique des Hautes Études, France; Université de Montréal, Canada · lea.maronet@huma-num.fr
Adam Faci
Centre National de la Recherche Scientifique (CNRS), France · adam.faci@huma-num.fr
Marion Charpier
École Nationale des Chartes, France · marion.charpier@chartes.psl.eu
Julien Schuh
Université Paris Nanterre, France · jschuh@parisnanterre.fr
Christopher Kermorvant
TEKLIA, France · kermorvant@teklia.com

This paper presents an ongoing collaboration between the French national infrastructure for the humanities and social sciences Huma-Num (https://www.huma-num.fr/), the Artificial Intelligence (AI) company Teklia (https://www.teklia.com/en), and the consortium pictorIA (https://pictoria.hypotheses.org/) to co-design and deploy an annotation infrastructure tailored to small, interpretive datasets in visual and audiovisual humanities. Our focus here is specifically on visual and audiovisual corpora where established annotation frameworks require significant adaptation. It reports on the early stages of a shared Arkindex instance, a web platform developed by Teklia, hosted on Huma-Num’s GPU infrastructure and on a series of methodological experiments around annotation practices. Our overarching question is how to move from exploratory experiments with AI to sustained, project-level production of high-quality, small-scale data that can facilitate interpretation and inform model training in specialized domains.

Rather than treating data as input for pattern-finding, our approach starts from the premise that annotations are historical and social artefacts: they encode the conceptual vocabularies, disciplinary norms, and power relations that shape what is visible or thinkable in a corpus. In the context of visual culture and heritage, this means that bounding boxes and labels are never neutral technical primitives, they crystallise particular ways of seeing, stabilising what counts as an “object,” where its boundaries lie, and which distinctions matter. At the same time, they can be limiting, especially when inherited from generic computer-vision pipelines that are poorly aligned with humanistic questions (Villaespesa / Murphy 2021). Our work therefore treats annotation formats and schemes as objects of research in their own right, not merely as intermediate steps between a corpus and aggregate statistics in a “distant viewing” logic. We deliberately explore alternative or complementary modes of annotation that can better capture ambiguity, composition, and inter-image relations. By doing so, we seek to expand the repertoire of what “counts” as annotation in AI-supported workflows, and to keep open a space for interpretive, small-scale work that is not immediately reduced to a fixed set of categories optimised for model training.

Our work therefore combines three dimensions:

On the technical side, Teklia’s Arkindex platform offers a modular environment in which heterogeneous tools (OCR/HTR engines, object detectors such as YOLO, similarity models such as CLIP, or custom models) can be orchestrated around document collections. The platform was built through sustained dialogue with researchers, prioritising fit-for-purpose design over generic versatility. Together with IR* Huma-Num, we have deployed a dedicated instance on public research infrastructure, with multi-user, multi-project, and fine-grained access control. This configuration is designed explicitly for small to medium sized, domain-specific corpora. The goal is not to maximise throughput, but to support iterative, collaborative annotation cycles where domain experts can intervene, correct, and reshape the data at each step (Maronet / Truc 2025).

Early usage statistics confirm a strong exploratory interest: tens of thousands of images have been ingested and hundreds of thousands of regions processed. Yet only a very small number of projects have, so far, exported data for downstream research or publication. This gap between experimentation and durable data production is the central problem our paper addresses. What is missing is a way to stabilise meaning across heterogeneous research questions. In the Humanities, the same object or image can support multiple, sometimes competing, readings and there is currently no shared or operationalised approach to record or compare these interpretations (Bönisch 2020). Similar challenges have been addressed for textual annotation in platforms such as CATMA (Gius et al. 2021) and INCEpTION (Klie et al. 2018), and for the sharing of ground-truth transcription data through initiatives like HTR-United (Chagué et al. 2022) -. However, comparable infrastructure for visual corpora remains largely absent: deciding what counts as a region, an object, or a meaningful relation is itself an interpretive act, one that cannot be resolved by borrowing the relatively stabilised conventions of textual markup. Visual annotation also involves spatial and compositional dimensions for which no shared formal vocabulary yet exists across disciplines. We argue that the bottleneck is the absence of shared repertoires of annotation practices that bridge the technical affordances of the platform and the interpretive needs of researchers. To tackle this, the pictorIA Annotation working group has organised thematic meetings to map existing tools and practices, explicitly distinguishing between structural annotation (locating regions of interest, segmenting pages, identifying objects) and semantic annotation (assigning categories, descriptions, relationships, or interpretive tags). We complemented this survey with practical workshops in which participants annotate the same images using three modes: directly in Arkindex, through remote collaborative whiteboards, and on paper with pens and post-its. The analogue exercises, in particular, have proved crucial to surface needs that are hard to express within the constraints of existing UIs: annotations about rhythm and composition, cross-image relationships in a corpus, or concerns about how annotations will be maintained and shared over time. From these experiments, several challenges emerge that speak directly to the DH2026 theme “Annotating | Beyond Patterns: Interpretation with Small Data”:

  • Designing for small data as a feature, not a limitation. For many humanities projects, the aim is not to annotate millions of images, but to deeply annotate hundreds. The platform must therefore privilege transparency, versioning, and ease of iterative refinement over pure scale. We discuss how Arkindex’s project structure, task configuration, and export mechanisms can be tuned to encourage small, self-contained datasets that remain intelligible years later.
  • Making annotation ontologies explicit and negotiable. Choices about labels, taxonomies, and relationships are not neutral. Working across heterogenous corpora (from the photographs of imperial and colonial armed conflicts of EyCon to the Japanese historical pictures of HikarIA) has made it especially clear that ontological decisions must be treated as first-class objects: they are documented, discussed across disciplines, and revised in light of community feedback. We present initial strategies to embed these ontologies within Arkindex in a way that remains visible to annotators and exportable for downstream analysis.
  • Connecting small, curated datasets to AI model training. One key motivation for building high-quality, small-scale annotation sets is to adapt models for under-resourced domains (e.g. specific historical typography, marginal visual genres, or non-hegemonic iconographies). We outline a workflow in which curated annotations produced by researchers feed back into Teklia’s model training pipelines, and show how even relatively small datasets can substantially improve performance on specialised tasks. This raises questions about how researchers can retain control over the use of their data in training, and how credit and governance should be organised.
  • Addressing algorithmic bias through situated annotation. Rather than assuming that bias lies only in the data in a statistical sense, we treat annotations as situated, accountable acts (Hovy / Prabhumoye 2021). We demonstrate how small, carefully documented datasets can be used both to reveal and to counteract algorithmic biases: for instance, by deliberately annotating under-represented iconographies, or by capturing ambiguity instead of forcing binary labels. The platform’s export features are crucial here, allowing researchers to recreate and critique the full genealogy of a dataset.

The contribution of this paper is threefold. Conceptually, it offers a framework for thinking about annotation in the age of AI as the construction of shared interpretive repertoires, not just the production of training labels. Methodologically, it describes a set of co-design practices that other communities can reuse when building annotation infrastructures for small data. Technically, it presents the first results and road-map of the Arkindex instance co-developed by Teklia, Huma-Num, and pictorIA, including planned developments such as IIIF integration, improved model-training ergonomics, and more guided export templates for humanities projects.

By shifting the focus from “big” pattern extraction to careful, small-scale annotation work, we argue that digital humanities can play a decisive role in shaping how AI systems are trained and evaluated in specialised cultural domains. Rather than passively adapting to off-the-shelf models, communities like pictorIA can build collectively governed annotation ecosystems in which data, tools, and interpretive frameworks are co-designed. Our presentation brings together three linked contributions: a conceptual framework for treating annotation as an interpretive and situated research practice; a set of annotation methods designed for small, heterogeneous visual and audiovisual corpora, which operationalise and test this framework; and practical case studies that evaluate these methods in real research contexts while reflecting on the conditions under which such practices can be supported, scaled, and partly industrialised through a shared annotation infrastructure. By following the process from theoretical framing to methodological experimentation and infrastructural implementation, the paper addresses a single overarching problem: how digital humanities communities can co-design annotation ecosystems in which data, tools, and interpretive frameworks remain collectively governed.

References
  1. Bönisch, Dominik. (2020): “The Curator’s Machine: Clustering of Museum Collection Data through Annotation of Hidden Connection Patterns between Artworks.”, in: International Journal for Digital Art History, no. 5. DOI: 10.11588/dah.2020.5.75953.
  2. Chagué, Alix / Clérice, Thibault / Romary, Laurent. (2022): “HTR-United : un écosystème pour une approche mutualisée de la transcription automatique des écritures manuscrites”, in: HAL. HAL:https://inria.hal.science/hal-04124743hal-04124743.
  3. Cordell, Ryan. (2020): “Machine Learning + Libraries: A Report on the State of the Field.”, in: Library of Congress. <https://labs.loc.gov/static/labs/work/reports/Cordell-LOC-ML-report.pdf> [16.05.2026].
  4. Gius, Evelyn / Meister, Jan C. / Meister, Malte / Petris, Marco / Gerstorfer, Dominik / Akazawa. Mari. (2024): CATMA. <https://zenodo.org/records/12092195> [16.05.2026].
  5. Hovy, Dirk / Prabhumoye, Shrimai. (2021): “Five Sources of Bias in Natural Language Processing”, in: Language and Linguistics Compass 15, no. 8 (2021): e12432. DOI: 10.1111/lnc3.12432.
  6. Klie, Jan-Christoph / Bugert, Michael / Boullosa, Beto / Eckart de Castilho, Richard / Gurevych, Iryna. (2018): “The INCEpTION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation.”, in: Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations, Santa Fe, New Mexico: 5–9. Association for Computational Linguistics. <https://aclanthology.org/C18-2002/> [16.05.2026].
  7. Maronet, Léa / Truc, Alice. (2025): “Improving Workflows in Digital Art History: Sharing Annotations for Cultural Heritage Image Segmentation and Object Detection.”, in: Transformations, A DARIAH Journal 1, June 2025. DOI: 10.46298/transformations.14591.
  8. Strien, Daniel van / Bell, Mark / McGregor, Nora R. / Trizna, Michael. (2022): “An Introduction to AI for GLAM”, in: Proceedings of the Second Teaching Machine Learning and Artificial Intelligence Workshop, PMLR, 2022 March 14: 20–24. <https://proceedings.mlr.press/v170/strien22a.html> [16.05.2026].
  9. Truc, Alice / Maronet, Léa. (2025): “Toward the Standardisation of Annotation Data for Patrimonial Images: A Case of Layout Detection for Art Magazines”, in: Digital Humanities in the Nordic and Baltic Countries Publications 7, no. 3, 2025. DOI: 10.5617/dhnbpub.12268.
  10. Villaespesa, Elena / Murphy, Oonagh. (2021): “This Is Not an Apple! Benefits and Challenges of Applying Computer Vision to Museum Collections”, in: Museum Management and Curatorship 36, no. 4, 2021: 362–83. DOI: 10.1080/09647775.2021.1873827.