Daejeon, July 27–31
Archival research in the digital age is frequently characterized by "tool fatigue", where scholars must navigate a disconnected archipelago of software solutions (O’Sullivan 2022). A historian might store scans in a commercial cloud, manage citations in a desktop manager, perform OCR in a proprietary "black box", and attempt visualizations in yet another isolated platform. This disjointed approach leads to data silos, loss of context, and significant friction when moving between individual stages of research. Furthermore, the reliance on proprietary ecosystems often compromises the long-term sustainability and reproducibility of research data, which represent one of the core values of the FAIR principles and DH communities.
This workshop addresses these systemic issues by demonstrating a unified technical workflow that is modular, open-source, and accessible. It moves beyond the demonstration of single tools to focus on the connectivity between them. The core idea of the proposed pipeline is interoperability: ensuring that data flows seamlessly from on-site scanning to final visualization via standard protocols (WebDAV, APIs) and open formats (Markdown, JSON, XML). The workshop advocates for a shift from ad-hoc tool adoption to a structured, engineered approach to humanities data (Tomaselli 2023).
The workshop is designed as a guided, practical walkthrough of the entire data processing lifecycle, utilizing existing open-source tools. It is divided into three integrated modules, each focusing on a critical stage of the pipeline. The main focus is on the analysis of a archival dataset samples mostly in the form of personal typewritten correspondence in English and Russian.
Tools: OwnCloud, Zotero, Logseq
The foundation of the pipeline is a robust storage and metadata strategy. We begin with OwnCloud – a secure, open-source alternative to commercial cloud storage and integrate Zotero not just as a citation manager, but as a central metadata hub. We will demonstrate how to link archival scans stored in OwnCloud directly to Zotero items. Crucially, we will cover the use of Zotero’s API and tagging systems (e.g., assigning VIAF IDs), emphasizing the importance of ensuring clean, standardized metadata at the source. Along with files and metadata organization we will work with Logseq, a privacy-first, open-source knowledge management tool based on Markdown. Participants will learn the basics of "Personal Knowledge Management" (PKM) system that can be integrated directly with Zotero. This ensures that the research process is documented as rigorously as the results, fostering reproducibility and facilitating the eventual writing process.
Tool: Arkindex
Once data is stored and cataloged, the next challenge is transforming static images into machine-readable text. This module introduces Arkindex, an open-source platform for document processing developed by Teklia. Participants will learn how to bulk upload scans from their storage to Arkindex and apply integrated OCR engines such as Tesseract. We will also discuss the HTR possibilities of Arkindex and the potential deployment of Visual Language Models (VLMs), such as Qwen3-VL. Additionally, we perform automatic named-entity recognition (NER) to identify names, dates, and locations within the processed text.
Tool: Nodegoat
This module focuses on the synthesis of data. We will use Nodegoat as the primary research environment where metadata from Zotero and full-text data from Arkindex converge. Participants will learn to build a custom data model (defining Objects like Letters, Persons, Places) and populate it automatically via API ingestion. We will demonstrate how to cross-reference these datasets to reveal hidden connections, visualizing correspondence networks and temporal patterns. We will specifically look at how to link full-text transcripts from Arkindex with the biographical data in Nodegoat to create a rich, navigable database for advanced graph visualizations.
The workshop combines theoretical framing with intensive practical application. Each module will open with a brief discussion of the "Lessons Learned" regarding the specific tool’s sustainability and accessibility, followed by a live demonstration and a "follow-along" implementation period. Participants will be provided with a "Pipeline Starter Kit" containing necessary links, configuration files, and sample datasets to ensure they can set up a working prototype on their own devices during the session.
By the end of this workshop, participants will be able to:
This workshop is intended for historians, archivists, and any DH researchers who deal with network and correspondence analysis based on archival documents and are looking to streamline their workflow. It is particularly relevant for those working with multilingual or non-Latin script collections (e.g., Cyrillic), though the methods are universal. Intermediate familiarity with basic data concepts (metadata, CSVs) is helpful. No coding experience is strictly required. Participants are required to bring their own laptop. Expected number of participants: 15–20.
Michal Racyn is a postdoc researcher at Masaryk University (Brno, Czech Republic). Serving as an ICT Specialist at the Center for Information Technologies and HUME Lab, Michal has extensive experience in deploying digital technologies within both academic and non-academic settings. He has held a Fulbright fellowship at Columbia University and CEEPUS fellowship at the University of Vienna. He is a recipient of the Marie Skłodowska-Curie Actions Global Postdoctoral Fellowship for the project COLD-WAR-EURASIA, which utilizes the proposed archival data processing pipeline for large-scale analysis of transnational intellectual networks, focusing on Eurasianism in the context of US Cold War Slavic Studies.
This workshop is supported by the CESNET Development Fund, z.s.p.o.