Daejeon, July 27–31
Small datasets have been recognised as valuable across STEM disciplines, where they can offer precision, interpretability, and practical advantages complementary to big data approaches (Ali et al., 2016; Goldkind et al., 2018; Faraway and Augustin, 2018; Hekler et al., 2019; Bagchi, 2021a; 2021b; Werner et al., 2023; Hackenberg et al., 2025). Digital Humanities often privilege large-scale datasets to uncover patterns, but this emphasis raises critical questions about scale (Schöch, 2013; Borgman, 2015; Berry et al., 2015): when is “big” data necessary, and what kinds of insights emerge when working with smaller, more granular datasets? This paper addresses these questions through an analysis of the Journal of Open Humanities Data (JOHD). We chose this journal as our data source because it is the only multi-disciplinary data journal dedicated to the humanities, hence it provides a readily available single source of peer-reviewed descriptions of datasets (data papers) across the span of humanities disciplines (McGillivray et al., 2026) Our familiarity with the journal – through our roles on its editorial team and board – also offers valuable insight into its scope, practices, and evolving contributions to data-driven humanities research.
Our study focusses on papers that describe small-scale datasets, which, in our framework, include fewer than 10,000 records. It examines the prevalence, characteristics, and methodological framing of these datasets, highlighting their potential for further interpretation and reflexive research practices. We also focus on their reuse potential, as small datasets often contain rich, well-contextualised, and carefully curated information that can be productively reinterpreted, combined, or repurposed in ways that large-scale datasets may obscure. Their manageable size and specificity make them particularly valuable for methodological experimentation, pedagogical use, and cross-disciplinary inquiry, thereby extending their impact beyond the scope of the original research.
We argue that small data is not a limitation but a methodological choice that foregrounds the cultural and historical dimensions of data production. Such datasets can mitigate algorithmic bias and inform the development of AI models for under-resourced languages and specialized domains. Some types of data are also inherently small: in fields such as Ancient World Studies, for instance, many datasets derive from fragmentary textual traditions, limited manuscript witnesses, or archaeological discoveries whose quantity and condition are constrained by historical survival rather than by contemporary data-collection practices (Farina et al., 2025). These forms of scarcity produce small datasets not by design, but by the nature of the material itself, and they demand interpretive approaches attentive to context, provenance, and uncertainty.
A similar dynamic appears in Linguistics, offering a useful disciplinary contrast. Some linguistic resources – such as prescriptive morphological, syntactic, semantic, or phonological rules – constitute a finite set of norms and are inherently small. By contrast, descriptive corpora that capture the distribution and variation of contemporary language use expand continuously and are typically required in large quantities for training models. Because these two resource types diverge in purpose, the distributions they represent – such as frequencies and collocational patterns – also differ (Park et al. 2019), leading to fundamentally distinct scale requirements. Linguistic resources associated with prescriptive rules (e.g., morphological, syntactic, semantic, and phonological rules) tend to be small, whereas corpora that reflect real usage can grow indefinitely and are typically required in large quantities for training large language models (LLMs). Moreover, high-quality small annotated corpora with particularly carefully controlled internal consistency can meaningfully support linguistic modeling; for example, Jung and Kwon (2011) demonstrate that such corpora enable reliable prediction of prosodic breaks.
By analyzing a corpus of JOHD data papers, we identify strategies for generating high-quality small datasets and explore how these practices challenge assumptions about scale, openness, and interpretability in Digital Humanities research (Risam and Edwards, 2017; Ciula et al., 2021). This work contributes to debates on data ethics and methodological diversity, offering insights into how small data can shape both scholarly interpretation and the future of AI-driven humanities research.