Daejeon, July 27–31
Tool directories are numerous and constitute a well-established genre within the Digital Humanities: ranging from DiRT to Bamboo and TAPoR (3.0)https://tapor.ca/ (Grant et al. 2020) to large EU initiatives such as the Social Sciences and Humanities Open Marketplacehttps://marketplace.sshopencloud.eu/ or consortia of the German National Research Data Infrastructure (NFDI). They approach the clear and tangible need of our research communities for an overview of computational tools through curatorial and hierarchical approaches to knowledge organization. Dombrowski (2021) has coined the term directory paradox to highlight the inherent tension between the goals of building comprehensive, up-to-date registries and the reality of doing so with unstable funding and mostly volunteer labour, resulting in data silos with proprietary infrastructures (data models, backends, and frontends). Consequently, these tool directories must be understood as having failed in their ambition to provide a comprehensive, representative, and up-to-date reflection of the available possibilities in computational research and digital scholarship. At best, they represent snapshots or documentation of historical practices. A focus on volatile presentation layers and custom infrastructure that require constant maintenance means that the data—themselves a result of scholarly work—remain largely inaccessible to computational approaches and further use.
We therefore adopt an approach inspired by minimal computing, making, and Open Science to address the need for (tool) registries. As DH practitioners, we need an infrastructure enabling us to sustainably document and share our evolving knowledge about computational tools for research: which tools exist, for what purpose can they be used, who has used them, what was their experience doing so, and where and how can I learn to apply them? Minimal computing entails aligning the question “what do we need?” with the answer to “what do we have at hand?” (cf. Gil / Ortega 2016; Risam / Gil 2022). Given our dependence on research funding project cycles, our options are severely constrained: we must build upon existing datasets and utilise open, free, and established software and infrastructures.
This poster presents our proposal for an open basic infrastructure for (tool) registries. At the core of our proposal lies Wikidata.https://wikidata.org/ Wikidata is the world’s largest knowledge graph and widely used for sharing and aligning data in the GLAM sector (see Zhao 2022; Fischer / Ohlig 2019). It is based on the open Wikibase software developed by Wikimedia, and a sister project of Wikipedia. Wikidata shares the governance structure for user-generated and user-curated content of the wider Wikiverse. Any individual can contribute to and maintain entries relevant to their specific research, while all data are version controlled with persistent URIs. Unlike many infrastructures in the Digital Humanities, multilingualism of both interfaces and data is a fundamental feature of the Wikiverse. Wikidata enables the iterative development of minimal data models, the maintenance of datasets, and their aggregation into curated collections within community projects. On the data level, Wikidata allows direct utilization as five-star Linked Open Data (Berners-Lee 2009) via SPARQL, APIs, and the web interface.
To establish a common reference model for distributed datasets, we propose a basic data model for DH tools, defining minimal bibliographic and technical properties while remaining compatible with existing data models (e.g. Christophersen et al. 2023). The basic data model and baseline data enable the reuse of entries within independently curated collections with their own extensions to the data model. For our DH tool registry, we opted for the TaDiRAH taxonomyhttps://vocabs.dariah.eu/tadirah/ (Borek et al. 2021) for classifying items and sub-setting the data set.
To address the volatility of user-generated data, we provide workflows and scripts to automatically save and release version-controlled data sets from Wikidata on Zenodo (Dresselhaus / Grallert 2024; Grallert / Dresselhaus 2024–).
Wikidata enables the implementation of the proposed approach without requiring any additional software. Alternatively, Wikidata could be used exclusively as an authority file and data provider for custom frontends, as exemplified by Scholia (Nielsen et al. 2017). Ultimately, our proposal addresses the sustainability of project funding by ensuring a continuous contribution of data to the Digital Commons (Wittel 2013) in the form of Wikidata throughout the project duration and enabling the continued reuse of these data beyond the project lifecycle.
The scientific value of our effort is twofold: on the one hand, we provide well-documented workflows, scripts, and data models for open registries and we built the nucleus of a tool registry for digital humanities to solve Dombrowski’s paradox, on the other.