DH 2026

Daejeon, July 27–31

Wed, July 2914:00–15:30S047106
Short Paper

Machina Emblematica: A project of scholarly curiosity, digital alchemy, and emblematic wonder.

Michela Vignoli
AIT Austrian Institute of Technology, Austria · michela.vignoli@ait.ac.at
Chiara Palladino
Durham University, United Kingdom · chiara.palladino@durham.ac.uk
Rainer Simon
Independent Developer, Austria · rainer@rainersimon.io

Introduction

Generative AI technologies such as Large Language Models (LLMs) and Vision-Language (VL) systems offer new possibilities to analyze multimodal meaning at scale. By integrating these technologies, retrieval-augmented generation (RAG) systems have the potential to support a multimodal turn in the Digital Humanities (Smits / Wevers 2023; Li et al. 2024). In this short paper we present Machina Emblematica, a prototype designed to showcase the potential of RAG systems for increasing the accessibility of multimodal historical sources (Fig. 1).

Our case study is Joachim Camerarius’ 1668 edition of the Symbola et Emblemata (Camerarius 1668), the most important Emblem book in early modern Europe, covering an extensive range of information about zoology, mineralogy, and botany in the form of (partly allegorical, partly scientific) images and explanations in Latin. Emblem books, one of the most successful and original forms of book production in early modern Europe, consisted of clearly distinguished images (emblems) and accompanying text: typically an epigram and a motto, but also, as in the case of Camerarius, a longer commentary with rich and diverse information about the specimen illustrated in the image. Such information included etymological discussions, physical characteristics and behavior, and finally moral, religious, or political associations. Camerarius used the emblematic genre to create an encyclopedic work of natural history, providing a visual and written compendium of the early modern view of nature as a signifier of deeper moral and philosophical insights (Ashworth 2004). It is precisely in this bi-modality that lies the immense didactic and cultural effectiveness of emblem books (Enenkel / Smith 2017, p. 2), but also their complexity and difficulty of interpretation (Daly 2014), as the interaction of images and text and the complex web of intertextual references make this genre very challenging to understand for modern readers.

Materials and Methods

Machina Emblematica experiments with an innovative way to access the multimodal content of this genre. The prototype enables users to explore the indexed texts and images of the digital edition of the Symbola provided by the Münchener Digitale Bibliothek. In total, the corpus consists of 838 pages containing titles, captions, full texts, and 401 emblems. The image-text pairs provide an excellent use case to demonstrate how embedding-based retrieval can be used for retrieving specific contents and intermedial relations in digitized historical sources (Pan et al. 2024). The prototype advances existing solutions from projects enhancing accessibility and exploration of historical images and texts e.g. Digital Camerarius (Palladino et al. 2025), Emblematica Online (Green et al. 2017), and ONiT Explorer (Vignoli et al. 2024). The original aspect of the Machina is the combination of vector-based search with an assistant that generates responses to user queries in English based on the original images and texts. This provides a responsive interface where the user can explore the complexities of the source in many different ways, going beyond keyword searches.

The Machina consists of a web-based frontend and a multimodal RAG backend. The search backend is implemented using Marqo to index the cleaned text transcriptions and the extracted images. The multimodal sources retrieved by the search backend are passed to the model as a basis for answering the entered user query.

  • The images are vectorized with an open-CLIP model.
  • The full texts are chunked and embedded with a Flax Sentence Embeddings v4 model.
  • For the reasoning module we opted for Qwen2.5, a state-of-the-art VL-model that can process both visual and textual information (Bai et al. 2025).

The reasoning module contextualizes the user query based on the chat history to (i) generate an appropriate database query and (ii) determine which content modality to prioritize. The adapted query is then forwarded to the search backend. Retrieved results are normalized and consolidated into (i) a natural-language context string and (ii) an array of image links. To address input token limitations, only the prioritized modality is provided to the VL-model. Based on the retrieved data, the module generates an answer that includes an English summary of the most relevant textual and visual content. Previewed digital resources include direct links to their original sources and are re-ranked according to their relevance to the VL-model’s output (see Fig. 2). Overall, this experimental research prototype offers a transparent, low-risk AI approach aligned with the requirements of the EU AI Act, Article 50 that targets emblem book enthusiasts who want to explore the historical textual and visual materials.

Results and Discussion

Machina Emblematica presents an original approach to the exploration of knowledge articulated through the complex interaction of image and text. First empirical experiments indicate that the Machina can provide eased access to and meaningful interaction with the Neo-Latin texts and emblematic contents. As such, it is especially useful to support the reading of works of natural history, but it could also be expanded to other types of literature, such as early exploration and travel accounts. Preliminary results show that the applied Qwen model is capable of correctly understanding, translating, and contextualizing the text and image material in many cases. However, the tool is an early prototype and does have some limitations.

First, the search backend only passes one modality of the retrieved contents to the VL-model. As a result, the answer generated by the model can at times be detached from the missing text or image context. This is particularly problematic as the captions to the emblems are notoriously extraneous from the content of the images, thus the section titles and full text context is at times necessary for formulating meaningful answers. Second, the retrieval of relevant sources could be improved, for example by including additional metadata into the indices (e.g. keywords, semantic enrichments). Third, we observed that the VL-model erroneously interprets image contents in some cases (e.g. it misidentifies depicted animals, contents, or scenes). A more thorough assessment of the generated answers with respect to their factuality and introduced inaccuracies due to model hallucinations will be included in a future publication.

Future research will focus on improving the prototype to provide more accurate and relevant responses. This includes experimenting with alternative VL-models for generating the answers; further enriching the metadata and indexes to improve retrieval results; and exploring more advanced system architectures (e.g. Keerthana et al. 2025; Lin et al. 2025).

Figure 1: Screenshot of the Machina Emblematica Logo (Landing Page)

Figure 2: Screenshot of a Machina Camerarius Chatbot Output (Excerpt)

Data Availability Statement

The tool can be accessed at https://machina.ait.ac.at/. The source code is published under an MIT license at https://github.com/rsimon/machina-emblematica.

Funding Acknowledgement

The project received funding from the BMBF joint project HERMES, the Furman Humanities Center, and the Department of Classics and Ancient History at Durham University.

References
  1. Ashworth, William B. (2004): “Natural History and the Emblematic World View”, in: Grasping the World. The Idea of the Museum, by Donald Preziosi and Claire Farago, eds. London: Routledge, pp. 144-158.
  2. Bai, Shuai / Chen, Kequin / Liu, Xuejing et al. (2025): Qwen2.5-VL Technical Report. arXiv. <https://arxiv.org/abs/2502.13923> [22.04.2026].
  3. Camerarius, Joachim (1668): Symbolorum et Emblematum. Centuria quatuor collecta. Moguntia: Bourgeat.
  4. Daly, Peter M. (2014): The Emblem in Early Modern Europe. Contributions to the Theory of the Emblem. Routledge, Taylor & Francis Group.
  5. Enenkel, Karl A. E. / Smith, Paul J. (2017): “Introduction: Emblems and the Natural World (ca. 1530-1700)”, in: Enenkel, Karl A. E. / Smith, Paul J. (eds.) Emblems and the Natural World. Intersections : Interdisciplinary Studies in Early Modern Culture, volume 50. Leiden, Boston: Brill, pp. 1-40.
  6. Green, Harriett / Cole, Timothy / Han, Myung-Ja (2017): “Enhancing User Access through the Emblematica Open Emblem Portal”, in: Hoepel, Ingrid / McKeown, Simon (eds.) Emblems and Impact Volume II: Von Zentrum und Peripherie der Emblematik. Newcastle upon Tyne: Cambridge Scholars Publishing, pp. 117-127.
  7. Keerthana, Murugaraj / Salima, Lamsiyah / During, Marten / Theobald, Martin (2025): “Topic-RAG for historical newspapers: Enhancing information retrieval in humanities research through topic-based retrieval-augmented generation”, in: Computational Humanities Research 1: e15.  DOI: 10.1017/chr.2025.10018.
  8. Li, Xiaoxi / Jin, Jiajie / Zhou, Yujia / Zhang, Yuyao / Zhang, Peitian / Zhu, Yutao / Dou, Zhicheng (2024): From Matching to Generation: A Survey on Generative Information Retrieval. arXiv. <http://arxiv.org/abs/2404.14851> [22.04.2026].
  9. Lin, Xiaoqiang / Ghosh, Aritra / Low, Bryan Kian Hsiang et al. (2025): REFRAG: Rethinking RAG based Decoding. arXiv. <http://arxiv.org/abs/2509.01092> [22.04.2026].
  10. Palladino, Chiara / Vignoli, Michela / Wilson, Kathryn (2025): “Digital Camerarius – Tracing the Classical origins of Pre-Linnean Science”, in: DH2025 Conference, Accessibility & Citizenship, July 14-18, 2025, Lisbon, Portugal.
  11. Pan, James Jie / Wang, Jianguo / Li, Guoliang (2024): “Survey of vector database management systems”, in: The VLDB Journal 33 (5): 1591–1615. DOI: 10.1007/s00778-024-00864-x.
  12. Smits, Thomas / Wevers, Melvin (2023): “A multimodal turn in Digital Humanities. Using contrastive machine learning models to explore, enrich, and analyze digital visual historical collections”, in: Digital Scholarship in the Humanities 38 (3): 1267–1280. DOI: 10.1093/llc/fqad008.
  13. Vignoli, Michela / Gruber, Doris / Seidl, Michael (2025): “Revolution or Evolution? AI-Driven Retrieval of Nature Representations in Historical Prints”, in: Digital Scholarship in the Humanities, fqaf082. DOI: 10.1093/llc/fqaf082.