Daejeon, July 27–31
This poster presents a resource created to support the teaching of Digital humanities to students in the humanities, particularly in fields such as Classics, Archaeology, History, and Linguistics. As Digital humanities methodologies continue to develop rapidly and as fully comprehensive, humanities-cantered introductions to data formats remain limited, this resource provides a structured, practical guide for students and instructors seeking to gain fundamental data literacy. It is designed to help learners develop the conceptual, ethical, and technical skills necessary to create, interpret, and publish research data in a consistent, interoperable, and sustainable manner.
The resource is a modular book comprising of more than twenty interconnected chapters that introduce the major data formats and modelling approaches used in the Digital humanities. It begins with three foundational modules that define data formats, explain data modelling as an interpretive practice, and introduce the ethical considerations inherent in structuring and representing cultural material, drawing on established work in DH data modelling (Flanders / Jannidis 2019). These modules help students understand that decisions about syntax, semantics, hierarchy, granularity, and uncertainty deeply influence the questions scholars can pose and the conclusions they can draw. They also introduce FAIR principles (Wilkinson et al. 2016), which serve as a guiding framework throughout the entire guide.
Following this conceptual grounding, the book turns to tabular data and spreadsheets, the formats most familiar to students and the entry point for most humanities datasets. These chapters introduce the basics of spreadsheets and CSV files, explore the structure and logic of tabular data, and explain why CSV remains a foundational format across digital scholarship. Central to this section is the application of tidy data principles as articulated by Hadley Wickham (Wickham 2014). Students learn that each variable must form a column, each observation must form a row, and each type of observational unit must form its own table. These principles, together with controlled vocabularies, consistent data types, and clear naming conventions, teach that datasets remain legible, analysable, and easily adaptable to new tools and workflows.
Building on these foundations, the guide introduces data cleaning and validation techniques, emphasizing the importance of consistent spelling, standardized entries, the handling of uncertain or incomplete data, and the clear documentation of all decisions made during the cleaning process. Students then learn how to analyse their datasets using spreadsheet tools such as sorting, filtering, conditional formatting, and pivot tables, gaining experience in identifying patterns and linking quantitative summaries to humanistic interpretation. This first major segment concludes with modules on exporting, licensing, publishing, and sustaining datasets, including best practices for metadata creation, README documentation, licensing choices, and long-term stewardship (Gilliland 2016).
The second major section introduces Geographic Information Systems (GIS) in a way that is accessible for humanities students. It explains spatial data fundamentals, coordinate systems, gazetteers, and the principles of mapping humanities evidence. Students practice linking tabular data to geographic coordinates, explore web-based mapping tools, and learn how to create spatial datasets that can integrate with external resources such as Pleiades, GeoNames, Wikidata, and OpenStreetMap, thereby introducing core concepts of Linked Open Data. The emphasis is on understanding spatial relationships in cultural, historical, and archaeological contexts while maintaining methodological awareness.
After the introduction to spatial data, the guide shifts to non-relational formats, beginning with XML. These modules teach XML syntax, validation, schema creation, the design of document models for humanities corpora (Fallside / Walmsley 2004), and techniques for querying and transforming XML through XPath and XSLT. The section culminates in an introduction to XML database management with eXist-db (Meier 2022), giving students exposure to real-world workflows used in large-scale digital editions and text-based research infrastructures. The book intentionally focuses on XML fundamentals rather than TEI-specific guidelines, ensuring that students first understand the underlying principles before engaging with more specialized frameworks.
The final chapters introduce JSON, outlining its syntax, structural logic, and common use cases in digital scholarship. Students compare JSON to XML and CSV, learning how different formats serve different representational needs. The guide concludes with a module on converting between CSV, XML, and JSON while minimizing information loss, a crucial skill for interoperability in digital humanities projects. The resource is designed for both classroom use and self-study. Each module includes clear explanations, historical context, practical exercises, reflection questions, case studies, and curated recommendations for further reading. To support hands-on learning, the guide includes 35 datasets across domains such as inscriptions, manuscripts, texts, people, and places, allowing students to practice all stages of the data lifecycle. These components aim to provide a comprehensive, flexible, and humanities-centered introduction to data modelling and formats, preparing students for deeper engagement with digital methods while grounding them in best practices for ethical and sustainable scholarship.