Daejeon, July 27–31
Large-scale collections of cultural heritage data have long been in demand for computational humanities research. Access to them, however, remains limited. A number of factors can explain this: legal restrictions tie access to institutional or national affiliations; different standards for data representation and variable data quality limit their utility for users.
In recent years, research under the umbrella term “computational humanities” has successfully opened up cultural heritage collections for specific humanities research questions, by adopting a variety of methods from fields such as statistics and natural language processing (Karsdorp et al. 2021; “Computational Humanities Research” 2025; “Journal of Cultural Analytics” 2024) To better support researchers and their evolving research practices, cultural heritage institutions are in parallel actively experimenting with data labs to offer access to tools and datasets (Candela et al. 2022).
Against this background, the proposed workshop will explore novel opportunities for the analysis and exploration of digital cultural heritage collections which arise through 1) programmatic access to large multilingual and multimodal data collections, 2) the availability of large language and vision models such as LLama or OpenCLIP) for their processing and enrichment, and 3) semantic linking of such data across languages and modalities based on embeddings and vector representations.
The proposed workshop is organised by the interdisciplinary research project Impresso Media Monitoring of the Past - Beyond Borders (https://impresso-project.ch/) which leverages an unprecedented corpus of newspaper and radio archives and uses machine learning to pursue a paradigm shift in the processing, semantic enrichment, representation, exploration and study of historical media across modalities, time, languages, and national borders. The instructors are researchers in Computational Humanities and Digital History working in the Impresso project.
During the workshop, participants will experiment with different corpus creation methodologies and search methods using keywords, images captions, metadata, semantic enrichments and embeddings to retrieve documents from large-scale newspapers collections (articles, advertisements, images, etc.).
Some of the questions to be explored in this session are:
These experiments will involve reusing and adapting existing notebooks as well as creating new ones tailored to the objectives of the group.
The Impresso Datalab (https://impresso-project.ch/datalab/) is a platform that provides programmatic access to a growing corpus of historical newspaper and radio collections, models for the semantic indexation of external research documents (e. g. for named entity recognition, press agency detection), and finally the ability to link and align external research data to the Impresso corpus on the basis of embeddings and vector representations.
It enables custom analyses of the Impresso corpus and the semantic indexation of external document collections also with the help of models created by the project. Access to data and models is facilitated via the Impresso Public API, a dedicated Python library and via HuggingFace (https://huggingface.co/impresso-project). To keep the learning curve manageable, the Datalab includes Jupyter notebooks for documentation and templates to kickstart computational analysis, visualisation and enrichment. Its design and development is driven by five objectives:
Overall, the Datalab supports researchers who are confronted with digitised collections with variable depth of annotation and seek to overcome the limitations of off-the-shelf graphical user interfaces. The Impresso project works with a consortium of partnering libraries and archives on the grounds of legal agreements which respect diverse copyright regulations.
The outputs of the workshop will include first of all case-study specific notebooks and lessons learned for the use of embeddings for navigating the Impresso corpus. Beyond these concrete outputs, the workshop will offer experience-based insights into a number of open questions: How can cultural heritage data labs best support researchers in the age of LLMs and generative AI tools? What research scenarios can be realised with vector representations of diverse research datasets? To what extent do such approaches support humanities research workflows commonly characterised by the synthesis of diverse data sources? Finally, to what extent does the Impresso Datalab effectively support such research workflows? What are the remaining gaps and what are future prospects for filling them?
The results from the workshop will be documented at the group-level in the form of notebooks and GitHub issues to keep track of problems and ideas for enhancements, and at the level of individual participants through informal exchanges during the event and a post-event survey.
The event will be structured as follows:
09:30-10:00: Introduction to Impresso Web App & Datalab
10:30-11:00: Presentation of Research cases and splitting in groups
11:00-11:15: Short break
11:15-12:30: Working through Research cases I
12:30-13:30: Lunch
13.30-14.00: Intermediate results presentation
14.00-16.30: Working through Research cases II
16.30-17.00: Group presentation of the research case results
The workshop is open to a maximum of 20 participants, at least basic programming skills in Python are required; all participants should bring a laptop.
Candela, Gustavo / Sáez, María Dolores / Escobar Esteban, MPilar/ Marco-Such, Manuel (2022). “Reusing Digital Collections from GLAM Institutions.”Journal of Information Science48 (2): 2. https://doi.org/10.1177/0165551520950246.
“Computational Humanities Research.” 2025. https://2025.computational-humanities-research.org/.
“Journal of Cultural Analytics.” 2024. July 23. https://culturalanalytics.org/.