DH 2026

Daejeon, July 27–31

Fri, July 3114:00–15:30S106204-205
Short Paper

Small Data, Big Questions: Insights from the Journal of Open Humanities Data

Youngim Jung
Korea Institute of Science and Technology Information; University of Science and Technology · acorn@kisti.re.kr
Andrea Farina
King’s College London · andrea.farina@kcl.ac.uk
Barbara McGillivray
King’s College London · barbara.mcgillivray@kcl.ac.uk

Small datasets have been recognised as valuable across STEM disciplines, where they can offer precision, interpretability, and practical advantages complementary to big data approaches (Ali et al., 2016; Goldkind et al., 2018; Faraway and Augustin, 2018; Hekler et al., 2019; Bagchi, 2021a; 2021b; Werner et al., 2023; Hackenberg et al., 2025). Digital Humanities often privilege large-scale datasets to uncover patterns, but this emphasis raises critical questions about scale (Schöch, 2013; Borgman, 2015; Berry et al., 2015): when is “big” data necessary, and what kinds of insights emerge when working with smaller, more granular datasets? This paper addresses these questions through an analysis of the Journal of Open Humanities Data (JOHD). We chose this journal as our data source because it is the only multi-disciplinary data journal dedicated to the humanities, hence it provides a readily available single source of peer-reviewed descriptions of datasets (data papers) across the span of humanities disciplines (McGillivray et al., 2026) Our familiarity with the journal – through our roles on its editorial team and board – also offers valuable insight into its scope, practices, and evolving contributions to data-driven humanities research.

Our study focusses on papers that describe small-scale datasets, which, in our framework, include fewer than 10,000 records. It examines the prevalence, characteristics, and methodological framing of these datasets, highlighting their potential for further interpretation and reflexive research practices. We also focus on their reuse potential, as small datasets often contain rich, well-contextualised, and carefully curated information that can be productively reinterpreted, combined, or repurposed in ways that large-scale datasets may obscure. Their manageable size and specificity make them particularly valuable for methodological experimentation, pedagogical use, and cross-disciplinary inquiry, thereby extending their impact beyond the scope of the original research.

We argue that small data is not a limitation but a methodological choice that foregrounds the cultural and historical dimensions of data production. Such datasets can mitigate algorithmic bias and inform the development of AI models for under-resourced languages and specialized domains. Some types of data are also inherently small: in fields such as Ancient World Studies, for instance, many datasets derive from fragmentary textual traditions, limited manuscript witnesses, or archaeological discoveries whose quantity and condition are constrained by historical survival rather than by contemporary data-collection practices (Farina et al., 2025). These forms of scarcity produce small datasets not by design, but by the nature of the material itself, and they demand interpretive approaches attentive to context, provenance, and uncertainty.

A similar dynamic appears in Linguistics, offering a useful disciplinary contrast. Some linguistic resources – such as prescriptive morphological, syntactic, semantic, or phonological rules – constitute a finite set of norms and are inherently small. By contrast, descriptive corpora that capture the distribution and variation of contemporary language use expand continuously and are typically required in large quantities for training models. Because these two resource types diverge in purpose, the distributions they represent – such as frequencies and collocational patterns – also differ (Park et al. 2019), leading to fundamentally distinct scale requirements. Linguistic resources associated with prescriptive rules (e.g., morphological, syntactic, semantic, and phonological rules) tend to be small, whereas corpora that reflect real usage can grow indefinitely and are typically required in large quantities for training large language models (LLMs). Moreover, high-quality small annotated corpora with particularly carefully controlled internal consistency can meaningfully support linguistic modeling; for example, Jung and Kwon (2011) demonstrate that such corpora enable reliable prediction of prosodic breaks.

By analyzing a corpus of JOHD data papers, we identify strategies for generating high-quality small datasets and explore how these practices challenge assumptions about scale, openness, and interpretability in Digital Humanities research (Risam and Edwards, 2017; Ciula et al., 2021). This work contributes to debates on data ethics and methodological diversity, offering insights into how small data can shape both scholarly interpretation and the future of AI-driven humanities research.

References
  1. Ali, Mohammad, Puput Ichwatus Sholihah, Kawsar Ahmed, and Sri Palupi Prabandari. 2016. Small Data and Big Data: Combination make better Decision. International Journal of Research in Management, Economics and Commerce, 6(10), 1–6
  2. Bagchi, Mayukh. 2021a. A large scale, knowledge intensive domain development methodology. Knowledge Organization 48 (1):8-23.
  3. Bagchi, Mayukh. 2021b. Towards Knowledge Organization Ecosystem (KOE). Cataloging & Classification Quarterly, 59 (8): 740–56. doi:10.1080/01639374.2021.1998282.
  4. Berry, David M., Erik Borra, Anne Helmond, Jean-Christophe Plantin, and Jill Walker Rettberg. 2015. 'The Data Sprint Approach: Exploring the field of Digital Humanities through Amazon’s Application Programming Interface', Digital Humanities Quarterly, vol. 9, no. 3. <http://www.digitalhumanities.org/dhq/vol/9/3/000222/000222.html>
  5. Borgman, Christine L. 2016. Big Data, Little Data, No Data: Scholarship in the Networked World. First MIT Press paperback edition. The MIT Press.
  6. Ciula, Arianna, Miguel Vieira, Ginestra Ferraro, et al. 2021. “Small Data and Process in Data Visualization: The Radical Translations Case Study.” Version 1. Preprint, arXiv. https://doi.org/10.48550/ARXIV.2110.09349.
  7. Faraway, Julian J., and Nicole H. Augustin. 2018. “When Small Data Beats Big Data.” Statistics & Probability Letters 136 (May): 142–45. https://doi.org/10.1016/j.spl.2018.02.031.
  8. Farina, Andrea, Marongiu, Paola, Bru, Mathilde and Borkowski, Daniele. "When Data Meets the Past: Data Collection, Sharing, and Reuse in Ancient World Studies" Open Information Science, vol. 9, no. 1, 2025, pp. 20250014. https://doi.org/10.1515/opis-2025-0014
  9. Goldkind, Lauri, Mamello Thinyane, and Moon Choi. 2018. “Small Data, Big Justice: The Intersection of Data Science, Social Good, and Social Services.” Journal of Technology in Human Services 36 (4): 175–78. doi:10.1080/15228835.2018.1539369.
  10. Hackenberg, Maren, Sophia G. Connor, Fabian Kabus, June Brawner, Ella Markham, Mahi Hardalupas, Areeq Chowdhury, Rolf Backofen, Anna Köttgen, Angelika Rohde, Nadine Binder and Harald Binder. “Small Data Explainer - The impact of small data methods in everyday life.” ArXiv abs/2507.11773 (2025)
  11. Hekler EB, Klasnja P, Chevance G, Golaszewski NM, Lewis D, Sim I. Why we need a small data paradigm. BMC Med. 2019 Jul 17;17(1):133. doi: 10.1186/s12916-019-1366-x. PMID: 31311528; PMCID: PMC6636023.
  12. McGillivray et al. (2026). Open Data Adoption Across the Humanities: Insights from the Journal of Open Humanities Data. In Xiaoguang Wang, Marcia Lei Zeng, Jin Gao, Ke Zhao(Eds.) AI and Smart Data for Cultural Heritage. Edited by. Routledge.
  13. Park, Chongwon, Elizabethada Wright, David Beard, Ron Regal. 2019. Rethinking the teaching of grammar from the perspective of corpus linguistics. Linguistic Research 36(1), 35–65.
  14. Risam, Roopika, and Susan Edwards. 2017. Micro DH: Digital Humanities at the Small Scale. Digital Humanities 2017.
  15. Schöch, Christof. 2013 Big? Smart? Clean? Messy? Data in the Humanities. Journal of Digital Humanities, 2 (3), pp.2-13. ⟨hal-00920254⟩
  16. Werner, Jonas, Philipp Beisswanger, Christoph Schürger, Marco Klaiber, and Andreas Theissler. 2023. “From Data to Wisdom: A Review of Applications and Data Value in the Context of Small Data.” Procedia Computer Science 225: 1251–60. https://doi.org/10.1016/j.procs.2023.10.113.
  17. Youngim Jung and Hyuk-Chul Kwon. 2011. Consistency Maintenance in Prosodic Labeling for Reliable Prediction of Prosodic Breaks. In Proceedings of the 5th Linguistic Annotation Workshop, pages 38–46, Portland, Oregon, USA. Association for Computational Linguistics.