DH 2026

Daejeon, July 27–31

Short Paper

Escaping the Trap of Style: Using Reasoning-VLM for VisualCorrespondence Discovery

Tsz-Kin Chau
EPFL, Switzerland · tszkin.chau@epfl.ch
Xin Huang
EPFL, Switzerland · xin.huang@epfl.ch
Sarah Kenderdine
EPFL, Switzerland · sarah.kenderdine@epfl.ch

Introduction

Inquiry into influence, transmission, and iconographic continuity begins from similarity as a lead for investigation. We characterize this primitive as visual correspondence: a supercategory encompassing the observation of shared features between two distinct items, prior to the more committed relations the humanities recognize — such as source, copy, quotation, study, parody, witness, and influence — each of which encodes a causal claim that requires interpretation. Tracing the lineage of visual correspondence across vast visual culture supports influence studies, network reconstruction, archival serendipity, and scholarly editing.

From a computational perspective, we formalize correspondence discovery as image-to-image, top-n similarity retrieval. Existing embedding-based methods serve this end poorly in cross-medium heritage material: similarity collapses onto medium and execution rather than depicted content, because the dominant signal in the learned feature space is rarely the iconographic one. We name this failure mode the Trap of Style, using “style” as the working computational label for everything in the visual signal that is not the depicted content, in the lineage of content/style decompositions in computer vision (Gatys et al. 2016).(Gatys et al. 2016).

We propose the Interpretive Sampling Function, an instrument that lets the scholar choose which dimension of the polyvalent similarity space to sample, recasting retrieval as a means of hypothesis generation. We demonstrate the approach by tracing iconographic correspondences from the Panorama of the Battle of Murten (1893; hereafter the Murten Panorama) back to 15th- and 16th-century Swiss illustrated chronicles.

The Challenge of Polyvalent Similarity

While cultural heritage institutions (GLAMs) continue to digitize holdings at ever-increasing volume and resolution, the fundamental bottleneck in the age of digital abundance is no longer access, but discovery. Detecting correspondences between these digital objects is non-trivial due to the polyvalent nature of the visual signal. It operates across a spectrum ranging from low-level morphological features (colour, texture, and shape), through mid-level features (composition, local patterns, and regularities of execution), to high-level semantic concepts (iconography, narrative, and action).

Previous computational paradigms have struggled to navigate this spectrum. Early descriptor-based methods, such as Bag of Visual Words (BoVW) (Bergel et al. 2013), excelled at instance-level retrieval but failed across heterogeneous artistic styles. Style transfer approaches (Bergel et al. 2013), excelled at instance-level retrieval but failed across heterogeneous artistic styles. Style transfer approaches (Ufer et al. 2021) offered a solution but they required training a neural network. Multimodal contrastive models (Ufer et al. 2021) offered a solution but they required training a neural network. Multimodal contrastive models (Radford et al. 2021) promise a universal semantic bridge, yet the Trap of Style persists: when applied to historical visual culture, they consistently cluster by medium and execution rather than by depicted content, rendering cross-temporal iconographic discovery impractical.(Radford et al. 2021) promise a universal semantic bridge, yet the Trap of Style persists: when applied to historical visual culture, they consistently cluster by medium and execution rather than by depicted content, rendering cross-temporal iconographic discovery impractical.

Methodology: Expert-Guided Semantic Embedding

We propose Prompt-Guided Semantic Embedding: instead of relying on direct visual embeddings, we introduce an intermediate layer of Reasoning Vision-Language Models (VLMs), leveraging their abstract visual understanding capabilities. By instructing a Reasoning VLM (Gemini 2.5 Pro) to ignore medium and execution and describe only narrative action, we extract representational and environmental natural-language descriptions (Figure 4). These descriptions are then embedded into vector space, creating an index aligned with scholarly intent rather than medium and execution.

Case Study: Visual Correspondence of the Murten Panorama

We validated this approach with a case study of Murten Panorama, an event-rich battle painting composed of numerous subscenes. The retrieval task was to track its iconographic inheritance back to its sources in the 15th- and 16th-century Swiss Illustrated Chronicles

Digitized and available at https://www.e-codices.unifr.ch/

. To establish a ground truth, we collaborated with a medievalist to identify six known visual correspondences. (Figure 1).

Figure Recurring appearance of the same visual narrative “Traveling Women” in the 15th- and 16th-century Swiss Illustrated Chronicles and the Murten Panorama

Experimental Results

Embedding Experiment

We compared our method against four baseline embedding models, each representing a distinct paradigm. Results were visualized using UMAP (Figure 2).

Baselines. The embedding spaces cluster by manuscript: medium and execution dominate the similarity signal.

Figure UMAP visualization of the four baseline embedding models, coloured by manuscript. Ordered from top-left to bottom-right: ResNet-50 (Visual Feature), SigLip (Multimodal), DinoV2 (Dense Feature), and DreamSim (Similarity Dataset Finetuned)

Our method. The Reasoning-VLM pipeline collapses these boundaries. Manuscripts mix freely while semantic clustering emerges as the dominant structure, indicating that the descriptions captured semantic content over medium and execution.

Figure UMAP visualization of our method. Left to right: coloured by manuscript, coloured by LDA topics learned from the generated descriptions (n=9)
Figure UMAP visualization of our method. Left to right: coloured by manuscript, coloured by LDA topics learned from the generated descriptions (n=9) — Wordcloud of representational (left) and environmental (right) descriptions

Retrieval Experiment

We employ a symmetric retrieval strategy: the query image is processed through the same VLM pipeline used for indexing, and the embeddings of the representational and environmental descriptions are fused with a user-defined weight.

Retrieval Performance: Our method successfully retrieved the ground truth correspondence within the top-5 results in 50% of test cases (3 out of 6) (Figure 5). In contrast, the SigLIP baseline was unable to retrieve the ground truth within the top-5 for any of the cases. Notably, SigLIP failed to retrieve any candidates from the monochrome manuscript.

Figure Comparison of Top-5 result of the 3 successful test cases between our method and SigLIP

Related Work

Recent research in visual discovery for cultural heritage has increasingly leveraged contrastive learning models for similarity retrieval (Springstein et al. 2021; Offert / Bell 2023; Sridhar et al. 2026). A growing cluster of projects applies CLIP-family multimodal embeddings to cultural heritage collections for natural-language and reverse-image search (Springstein et al. 2021; Offert / Bell 2023; Sridhar et al. 2026). A growing cluster of projects applies CLIP-family multimodal embeddings to cultural heritage collections for natural-language and reverse-image search (Smits / Wevers 2023; Roald et al. 2024; Mahowald / Lee 2024; Huang / Lee 2025), demonstrating the broad uptake of contrastive vision-language models in this space. Conversely, Arnold and Tilton demonstrated that textual embeddings from VLM captions can enhance image-to-image recommendation in a case study of a documentary photograph archive (Smits / Wevers 2023; Roald et al. 2024; Mahowald / Lee 2024; Huang / Lee 2025), demonstrating the broad uptake of contrastive vision-language models in this space. Conversely, Arnold and Tilton demonstrated that textual embeddings from VLM captions can enhance image-to-image recommendation in a case study of a documentary photograph archive (Arnold / Tilton 2024). While their work focuses on a single medium, ours focuses on cross-medium similarity retrieval. For art-historical captioning, prior work has explored ICONCLASS-grounded neural captioning (Arnold / Tilton 2024). While their work focuses on a single medium, ours focuses on cross-medium similarity retrieval. For art-historical captioning, prior work has explored ICONCLASS-grounded neural captioning (Cetinic 2021) and retrieval-augmented generation (Cetinic 2021) and retrieval-augmented generation (Wang et al. 2025). Semantic web ontologies offer alternative scaffolding for expert prompts (Wang et al. 2025). Semantic web ontologies offer alternative scaffolding for expert prompts (Carboni / Luca, de 2019; Sartini et al. 2023). Most directly related, Impett and Offert (2022) use paired textual concepts as conceptual axes to probe CLIP's “vector imaginary”. Both their work and ours treat scholar-specified queries as a hermeneutic tool on a vision-language embedding space (Carboni / Luca, de 2019; Sartini et al. 2023). Most directly related, Impett and Offert (2022) use paired textual concepts as conceptual axes to probe CLIP's “vector imaginary”. Both their work and ours treat scholar-specified queries as a hermeneutic tool on a vision-language embedding space (Impett / Offert 2022).(Impett / Offert 2022).

Discussion and Open Issue

Our results highlight a fundamental limitation: the lack of a benchmark dataset for evaluating polyvalent similarity. Such a resource, labelled with specific correspondence types, would enable the field to move from qualitative demonstrations to rigorous quantitative evaluation.

While commercial state-of-the-art VLMs present accessibility challenges, open-weight models make the approach increasingly sustainable. Since text embeddings cannot exhaust every dimension of polyvalent similarity, we envision an ensemble architecture similar to iArt, where experts dynamically combine various interpretive sampling functions, is the suitable expert-in-the-loop application architecture.