DH 2026

Daejeon, July 27–31

Fri, July 3114:00–15:30S013104
Long Paper

Human-Centered Explainability in Vision–Language Models: A Case Study in Art-historical Interpretation

Stefanie Schneider
Marburg University, Germany · stefanie.schneider@uni-marburg.de

Introduction

Recent advances in Vision–Language Models (VLMs) have greatly expanded the analytical repertoire of the digital humanities. By aligning visual and linguistic information in a shared embedding space, these models perform a wide range of tasks, from retrieval (Radford et al. 2021) and captioning (Li et al. 2023) to zero-shot classification (Zhai et al. 2022). Yet their versatility has also drawn sustained criticism: their internal mechanisms remain opaque, and the large, uncurated datasets on which they are trained reproduce ethical, sociotechnical, and epistemological biases (Bender et al. 2021; Bhalla et al. 2024; Birhane et al. 2021). For art-historical inquiry—where visual meaning is shaped by culturally sedimented conventions of style, iconography, and material practice—this opacity is especially consequential. Here, the problem of explainability becomes not merely technical but epistemic.

In this paper, we ask to what extent Explainable Artificial Intelligence (XAI) can render the visual logic of a specific VLM legible to human interpreters—and, by extension, how such legibility might contribute to a more critically robust methodology for digital art history. We focus on CLIP (Radford et al. 2021), a multimodal model that has been employed widely for art-historical retrieval tasks (Springstein et al. 2021; Offert / Bell 2023). To interrogate its epistemic affordances, we compare seven saliency-based methods across three paradigms: (1) gradient-based (Grad-CAM, Selvaraju et al. 2017; Grad-CAM++, Chattopadhyay et al. 2018; LayerCAM, Jiang et al. 2021; LeGrad, Bousselham et al. 2024); (2) score-based (ScoreCAM, Wang et al. 2020; gScoreCAM, Chen et al. 2022); and (3) CLIP-specific (CLIP Surgery, Li et al. 2025). Each produces, for a given text prompt, a saliency map indicating the relative contribution of an image region to the model’s predicted score. In an online survey with participants trained in art history, we assess the extent to which these computational visualizations align with human visual judgment. Our inquiry proceeds through three guiding questions: (1) How effectively do XAI methods localize iconographic motifs in artworks under zero-shot conditions? (2) To what extent do their visualizations correspond to expert perception? (3) Which conceptual or visual factors structure these correspondences? By examining where XAI methods succeed—or fail—in making CLIP’s internal mechanisms visible, we contribute to a broader methodological debate about how digital art history might critically engage with the epistemic structures of machine vision.

Background

Traditional XAI approaches are often motivated by the demand for transparency in decision-making systems, emphasizing their reproducibility and accountability (for overviews, see, e.g., Saeed / Omlin 2023; Hassija et al. 2024). Yet, as Miller (2019) argues, explanations are not neutral descriptions of causal mechanisms but inherently dialogical: they are contrastive ( Why this rather than that?), selective ( Which of the many possible causes should be prioritized?), and social ( How should this explanation be shaped for the person to whom it is addressed?). For art-historical inquiry—where interpretation emerges through negotiation across visual, historical, and cultural dimensions—this ‘social model of explanation’ is indispensable. Saliency-based visualization methods attempt to address this issue by generating spatially localized heatmaps that indicate which image regions contribute most strongly to a model's decision (see Figure 3 and Figure 4 for examples). However, it remains uncertain whether these methods truly disclose a model’s internal conceptual structure, or merely aestheticize its opacity. As Jain and Wallace (2019) have shown, attention mechanisms correlate only weakly with decision relevance. The present case study, therefore, shifts focus: rather than assuming interpretability as a property of the model, it examines interpretability as it is perceived by human viewers—asking how, and to what extent, computational explanations resonate with art-historical judgment.

Figure 1. Seven artworks selected for the online study, each paired with two target classes: Petrus Christus, A Goldsmith in his Shop (1449; a) with “convex mirror” and “girdle”; Franz von Stuck, Adam and Eve (c. 1920; b) with “arm outstretched” and “snake”; Antonello da Messina, Calvary (1475; c) with “John” and “thief”; Claude Monet, Japanese Footbridge (1899; d) with “bridge” and “flower”; Jean-Auguste-Dominique Ingres, Oedipus and the Sphinx (1808; e) with “left foot” and “Sphinx”; Sandro Botticelli, The Lamentation (c. 1490; f) with “sword” and “Virgin Mary”; Bartholomeus van der Helst, The Musician (1662; g) with “lustful” and “sheet music.”

Experimental Design

The study involved a within-subjects online survey, which was conducted using SoSci Survey between June and July 2025.

https://www.soscisurvey.de/, last accessed 26 April 2026.

Each participant viewed a series of image-class pairs, with two classes being considered for each of seven artworks (Figure 1). Participants first annotated regions of interest corresponding to each class, and then ranked the saliency maps generated by the aforementioned XAI methods according to how well the visualization corresponded to their own annotations. Incomplete rankings were imputed using the Multivariate Imputation by Chained Equations (MICE) algorithm (Buuren / Groothuis-Oudshoorn 2011) to ensure a complete dataset for within-subjects analysis. Inter-rater agreement was assessed using Kendall’s W, and differences between methods were evaluated through mean-rank comparisons.

A total of 33 participants were recruited from the student populations of the University of Munich and the University of Göttingen. Of these, 75.76% identified as female, 21.21% as male, and 3.03% as diverse. Ages ranged from 18 to 65 years (M = 42.21; SD = 19.08). Regarding art-historical expertise, 62.50% reported basic knowledge, 21.88% intermediate, and the remainder advanced proficiency. Although the sample size is only modest, a power analysis indicated sufficient sensitivity to detect medium effects (Kendall’s W 0.30) with an estimated statistical power of approximately 80 to 85%.

Figure 2. Evaluation results shown as divergent stacked bar charts, comparing seven visual explainability methods for each image and class. Kendall’s W is reported to assess inter-rater reliability.

Figure 3. Ground-truth annotations (a) and saliency maps for Antonello da Messina’s Calvary (1475), shown for the class “thief”: CLIP Surgery (b), LeGrad (c), ScoreCAM (d), gScoreCAM (e), Grad-CAM (f), Grad-CAM++ (g), and LayerCAM (h).

Figure 4. Ground-truth annotations (a) and saliency maps for Sandro Botticelli’s The Lamentation (c. 1490), shown for the class “Virgin Mary”: CLIP Surgery (b), LeGrad (c), ScoreCAM (d), gScoreCAM (e), Grad-CAM (f), Grad-CAM++ (g), and LayerCAM (h).

Results

Figure 2 shows divergent stacked bar charts illustrating the ranking distributions across image-class pairs. Three methods—CLIP Surgery, LeGrad, and ScoreCAM—consistently achieve the highest mean ranks, followed closely by gScoreCAM. In contrast, gradient-based methods (Grad-CAM, Grad-CAM++, and LayerCAM) are typically ranked lowest, indicating a weaker correspondence with human annotations. The overall consistency of the rankings suggests that participants have a broadly shared perceptual understanding of saliency coherence. However, Figure 2 also demonstrates that inter-rater reliability varies substantially by target class. Spatially bounded classes—such as the “snake” in von Stuck’s Adam and Eve or the “flower” in Monet’s Japanese Footbridge—exhibit high consensus (Kendall’s W > 0.7). Yet symbolic or interpretive categories—“lustful” or the “Virgin Mary”—yield dispersed rankings, reflecting the ambiguity in both human perception and model representation. Moreover, the misattribution of several figures in Botticelli’s Lamentation—where participants frequently labeled Mary Magdalene as the Virgin Mary—exemplifies the inherent instability of human interpretation, which depends on the domain knowledge of the viewer (Figure 4).

Discussion

These findings suggest that algorithmic explainability and human interpretation intersect only partially. Two structural insights emerge in particular. (1) Ambiguity of concepts. Interpretability declines as categories become more symbolic, contextual, or culturally loaded. The frequent misidentification of religious figures, such as the Virgin Mary, highlights how both human and machine vision struggle when confronted with complex iconographic motifs. This leads to: (2) Limits of representation. Saliency-based methods cannot recover what the model itself fails to encode. In Antonello da Messina’s Calvary, for instance, the class “thief” produces diffuse activations across all methods, which indicates that CLIP does not encode the class as a transferable visual concept (Figure 3). The difficulty may partly stem from the unnatural posture of the figures, but—more fundamentally—it reflects the fragmentary nature of the training corpus (Bender et al. 2021; Birhane et al. 2021). Ambiguity, in this case, is not simply a problem of human perception, but a structural property of the machine’s representational logic.

More fundamentally, these results underscore that explainability does not equate to understanding. Saliency maps visualize statistical correlations within an embedding space, rather than the cultural or historical semantics that give rise to those correlations. Explanation, as Miller (2019) emphasizes, is dialogical: it constructs shared meaning between an explainer and an explainee. In art-historical research, this dialogue must confront the asymmetries between algorithmic inference and human interpretation, recognizing that interpretability depends as much on the viewer’s expectations and disciplinary frameworks as on the model’s design. The shared variability of algorithmic and human interpretation therefore calls for hybrid evaluation frameworks that treat explainability not as a ground truth but as an interpretive dialogue or negotiation.

Conclusions

This case study demonstrates the following: (1) architecture-specific methods, such as CLIP Surgery, produce the most perceptually coherent explanations; (2) interpretability depends on both visual concreteness and conceptual stability; and (3) the epistemic value of explainability lies not only in transparency, but also in critical reflexivity. By treating the outputs of XAI methods as interpretive mediations rather than proofs of machine “understanding,” this work reframes explainability as a humanistic practice—a means of interrogating how visual meaning is computed, represented, and rendered intelligible. In doing so, it contributes to a broader agenda in the digital humanities: to engage with algorithmic systems not merely as heuristic instruments, but as cultural interlocutors whose modes of vision both reveal and reproduce the imaginaries of their training data.

References
  1. Bender, Emily M. / Gebru, Timnit / McMillan-Major, Angelina / Shmitchell, Shmargaret (2021): “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”, in: FAccT '21: 2021 ACM Conference on Fairness, Accountability, and Transparency: 610–623. ACM. DOI: 10.1145/3442188.3445922.
  2. Bhalla, Usha / Oesterling, Alex / Srinivas, Suraj / Calmon, Flávio P. / Lakkaraju, Himabindu (2024): “Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)”, in: Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024. <http://papers.nips.cc/paper_files/paper/2024/hash/996bef37d8a638f37bdfcac2789e835d-Abstract-Conference.html>.
  3. Birhane, Abeba / Prabhu, Vinay Uday / Kahembwe, Emmanuel (2021): “Multimodal Datasets: Misogyny, Pornography, and Malignant Stereotypes”. arXiv:2110.01963.
  4. Bousselham, Walid / Boggust, Angie W. / Chaybouti, Sofian / Strobelt, Hendrik / Kuehne, Hilde (2024): “LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity”. arXiv:2404.03214.
  5. van Buuren, S. / Groothuis-Oudshoorn, K. (2011): “mice: Multivariate Imputation by Chained Equations in R”, in: Journal of Statistical Software 45, 3: 1–67. DOI: 10.18637/jss.v045.i03.
  6. Chattopadhyay, Aditya / Sarkar, Anirban / Howlader, Prantik / Balasubramanian, Vineeth N. (2018): “Grad-CAM++: Generalized Gradient-Based Visual Explanations for Deep Convolutional Networks”, in: 2018 IEEE Winter Conference on Applications of Computer Vision, WACV 2018: 839–847. IEEE Computer Society. DOI: 10.1109/WACV.2018.00097.
  7. Chen, Peijie / Li, Qi / Biaz, Saad / Bui, Trung / Nguyen, Anh (2022): “gScoreCAM: What Objects is CLIP Looking At?”, in: Computer Vision – ACCV 2022 – 16th Asian Conference on Computer Vision: 588–604. Springer. DOI: 10.1007/978-3-031-26316-3_35.
  8. Hassija, Vikas / Chamola, Vinay / Mahapatra, Atmesh / Singal, Abhinandan / Goel, Divyansh / Huang, Kaizhu / Scardapane, Simone / Spinelli, Indro / Mahmud, Mufti / Hussain, Amir (2024): “Interpreting Black-Box Models: A Review on Explainable Artificial Intelligence”, in: Cognitive Computation 16, 1: 45–74. DOI: 10.1007/S12559-023-10179-8.
  9. Jain, Sarthak / Wallace, Byron C. (2019): “Attention is not Explanation”, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019: 3543–3556. Association for Computational Linguistics. DOI: 10.18653/V1/N19-1357.
  10. Jiang, Peng-Tao / Zhang, Chang-Bin / Hou, Qibin / Cheng, Ming-Ming / Wei, Yunchao (2021): “LayerCAM: Exploring Hierarchical Class Activation Maps for Localization”, in: IEEE Transactions on Image Processing 30: 5875–5888. DOI: 10.1109/TIP.2021.3089943.
  11. Li, Junnan / Li, Dongxu / Savarese, Silvio / Hoi, Steven C. H. (2023): “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”, in: International Conference on Machine Learning, ICML 2023: 19730–19742. PMLR. <https://proceedings.mlr.press/v202/li23q.html>.
  12. Li, Yi / Wang, Hualiang / Duan, Yiqun / Zhang, Jiheng / Li, Xiaomeng (2025): “A Closer Look at the Explainability of Contrastive Language-image Pre-training”, in: Pattern Recognition 162: 111409. DOI: 10.1016/J.PATCOG.2025.111409.
  13. Miller, Tim (2019): “Explanation in Artificial Intelligence: Insights from the Social Sciences”, in: Artificial Intelligence 267: 1–38. DOI: 10.1016/J.ARTINT.2018.07.007.
  14. Offert, Fabian / Bell, Peter (2023): “imgs.ai: A Deep Visual Search Engine for Digital Art History”, in: International Conference of the Alliance of Digital Humanities Organizations, DH2022. DOI: 10.5281/zenodo.8107778.
  15. Radford, Alec / Kim, Jong Wook / Hallacy, Chris / Ramesh, Aditya / Goh, Gabriel / Agarwal, Sandhini / Sastry, Girish / Askell, Amanda / Mishkin, Pamela / Clark, Jack / Krueger, Gretchen / Sutskever, Ilya (2021): “Learning Transferable Visual Models From Natural Language Supervision”, in: Proceedings of the 38th International Conference on Machine Learning, ICML 2021: 8748–8763. PMLR. <http://proceedings.mlr.press/v139/radford21a.html>.
  16. Saeed, Waddah / Omlin, Christian W. (2023): “Explainable AI (XAI): A Systematic Meta-survey of Current Challenges and Future Opportunities”, in: Knowledge-Based Systems 263: 110273. DOI: 10.1016/J.KNOSYS.2023.110273.
  17. Selvaraju, Ramprasaath R. / Cogswell, Michael / Das, Abhishek / Vedantam, Ramakrishna / Parikh, Devi / Batra, Dhruv (2017): “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization”, in: IEEE International Conference on Computer Vision, ICCV 2017: 618–626. IEEE Computer Society. DOI: 10.1109/ICCV.2017.74.
  18. Springstein, Matthias / Schneider, Stefanie / Rahnama, Javad / Hüllermeier, Eyke / Kohle, Hubertus / Ewerth, Ralph (2021): “iART: A Search Engine for Art-historical Images to Support Research in the Humanities”, in: MM '21: ACM Multimedia Conference: 2801–2803. ACM. DOI: 10.1145/3474085.3478564.
  19. Wang, Haofan / Wang, Zifan / Du, Mengnan / Yang, Fan / Zhang, Zijian / Ding, Sirui / Mardziel, Piotr / Hu, Xia (2020): “Score-CAM: Score-weighted Visual Explanations for Convolutional Neural Networks”, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR Workshops 2020: 111–119. Computer Vision Foundation / IEEE. DOI: 10.1109/CVPRW50498.2020.00020.
  20. Zhai, Xiaohua / Wang, Xiao / Mustafa, Basil / Steiner, Andreas / Keysers, Daniel / Kolesnikov, Alexander / Beyer, Lucas (2022): “LiT: Zero-Shot Transfer with Locked-image text Tuning”, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022: 18102–18112. IEEE. DOI: 10.1109/CVPR52688.2022.01759.