Daejeon, July 27–31
This paper presents an ongoing collaboration between the French national infrastructure for the humanities and social sciences Huma-Num (https://www.huma-num.fr/), the Artificial Intelligence (AI) company Teklia (https://www.teklia.com/en), and the consortium pictorIA (https://pictoria.hypotheses.org/) to co-design and deploy an annotation infrastructure tailored to small, interpretive datasets in visual and audiovisual humanities. Our focus here is specifically on visual and audiovisual corpora where established annotation frameworks require significant adaptation. It reports on the early stages of a shared Arkindex instance, a web platform developed by Teklia, hosted on Huma-Num’s GPU infrastructure and on a series of methodological experiments around annotation practices. Our overarching question is how to move from exploratory experiments with AI to sustained, project-level production of high-quality, small-scale data that can facilitate interpretation and inform model training in specialized domains.
Rather than treating data as input for pattern-finding, our approach starts from the premise that annotations are historical and social artefacts: they encode the conceptual vocabularies, disciplinary norms, and power relations that shape what is visible or thinkable in a corpus. In the context of visual culture and heritage, this means that bounding boxes and labels are never neutral technical primitives, they crystallise particular ways of seeing, stabilising what counts as an “object,” where its boundaries lie, and which distinctions matter. At the same time, they can be limiting, especially when inherited from generic computer-vision pipelines that are poorly aligned with humanistic questions (Villaespesa / Murphy 2021). Our work therefore treats annotation formats and schemes as objects of research in their own right, not merely as intermediate steps between a corpus and aggregate statistics in a “distant viewing” logic. We deliberately explore alternative or complementary modes of annotation that can better capture ambiguity, composition, and inter-image relations. By doing so, we seek to expand the repertoire of what “counts” as annotation in AI-supported workflows, and to keep open a space for interpretive, small-scale work that is not immediately reduced to a fixed set of categories optimised for model training.
Our work therefore combines three dimensions:
On the technical side, Teklia’s Arkindex platform offers a modular environment in which heterogeneous tools (OCR/HTR engines, object detectors such as YOLO, similarity models such as CLIP, or custom models) can be orchestrated around document collections. The platform was built through sustained dialogue with researchers, prioritising fit-for-purpose design over generic versatility. Together with IR* Huma-Num, we have deployed a dedicated instance on public research infrastructure, with multi-user, multi-project, and fine-grained access control. This configuration is designed explicitly for small to medium sized, domain-specific corpora. The goal is not to maximise throughput, but to support iterative, collaborative annotation cycles where domain experts can intervene, correct, and reshape the data at each step (Maronet / Truc 2025).
Early usage statistics confirm a strong exploratory interest: tens of thousands of images have been ingested and hundreds of thousands of regions processed. Yet only a very small number of projects have, so far, exported data for downstream research or publication. This gap between experimentation and durable data production is the central problem our paper addresses. What is missing is a way to stabilise meaning across heterogeneous research questions. In the Humanities, the same object or image can support multiple, sometimes competing, readings and there is currently no shared or operationalised approach to record or compare these interpretations (Bönisch 2020). Similar challenges have been addressed for textual annotation in platforms such as CATMA (Gius et al. 2021) and INCEpTION (Klie et al. 2018), and for the sharing of ground-truth transcription data through initiatives like HTR-United (Chagué et al. 2022) -. However, comparable infrastructure for visual corpora remains largely absent: deciding what counts as a region, an object, or a meaningful relation is itself an interpretive act, one that cannot be resolved by borrowing the relatively stabilised conventions of textual markup. Visual annotation also involves spatial and compositional dimensions for which no shared formal vocabulary yet exists across disciplines. We argue that the bottleneck is the absence of shared repertoires of annotation practices that bridge the technical affordances of the platform and the interpretive needs of researchers. To tackle this, the pictorIA Annotation working group has organised thematic meetings to map existing tools and practices, explicitly distinguishing between structural annotation (locating regions of interest, segmenting pages, identifying objects) and semantic annotation (assigning categories, descriptions, relationships, or interpretive tags). We complemented this survey with practical workshops in which participants annotate the same images using three modes: directly in Arkindex, through remote collaborative whiteboards, and on paper with pens and post-its. The analogue exercises, in particular, have proved crucial to surface needs that are hard to express within the constraints of existing UIs: annotations about rhythm and composition, cross-image relationships in a corpus, or concerns about how annotations will be maintained and shared over time. From these experiments, several challenges emerge that speak directly to the DH2026 theme “Annotating | Beyond Patterns: Interpretation with Small Data”:
The contribution of this paper is threefold. Conceptually, it offers a framework for thinking about annotation in the age of AI as the construction of shared interpretive repertoires, not just the production of training labels. Methodologically, it describes a set of co-design practices that other communities can reuse when building annotation infrastructures for small data. Technically, it presents the first results and road-map of the Arkindex instance co-developed by Teklia, Huma-Num, and pictorIA, including planned developments such as IIIF integration, improved model-training ergonomics, and more guided export templates for humanities projects.
By shifting the focus from “big” pattern extraction to careful, small-scale annotation work, we argue that digital humanities can play a decisive role in shaping how AI systems are trained and evaluated in specialised cultural domains. Rather than passively adapting to off-the-shelf models, communities like pictorIA can build collectively governed annotation ecosystems in which data, tools, and interpretive frameworks are co-designed. Our presentation brings together three linked contributions: a conceptual framework for treating annotation as an interpretive and situated research practice; a set of annotation methods designed for small, heterogeneous visual and audiovisual corpora, which operationalise and test this framework; and practical case studies that evaluate these methods in real research contexts while reflecting on the conditions under which such practices can be supported, scaled, and partly industrialised through a shared annotation infrastructure. By following the process from theoretical framing to methodological experimentation and infrastructural implementation, the paper addresses a single overarching problem: how digital humanities communities can co-design annotation ecosystems in which data, tools, and interpretive frameworks remain collectively governed.