DH 2026

Daejeon, July 27–31

Workshop

Leveraging Emerging Datasets for Critical Research on Books and Reader Communities

Yuerong Hu
Luddy School of Informatics, Computing and Engineering, Indiana University Bloomington, United States of America · yuerhu@iu.edu
Cici Ling
Luddy School of Informatics, Computing and Engineering, Indiana University Bloomington, United States of America · ccling@iu.edu
Kai Li
School of Information Sciences, The University of Tennessee, Knoxville · kli16@utk.edu
Andrew Zalot
College of Education, East Carolina University · zalota25@ecu.edu
Wenyi Shang
School of Information Science & Learning Technologies, University of Missouri · wenyishang@missouri.edu
Micah Bateman
School of Library and Information Science, University of Iowa · micah-bateman@uiowa.edu
Peizhen Wu
English Department, University of Illinois Urbana-Champaign · peizhen4@illinois.edu
Natasha Johnson
Language Technologies Institute, Carnegie Mellon University · nmj@alumni.stanford.edu

DH-Workshop Topic and Introduction.

In the last two decades, unprecedented volumes of data about books and reader communities have been available online, such as user-added genre tags (Walsh / Antoniak 2021), “BookTok” videos (Murray 2024), fanfiction (Johnson et al. 2025), and web novels (Yu / Pianzola 2024). These internet datasets have been increasingly used for digital humanities (DH) research on reading, literature reception, and contemporary book culture (Murray 2018). Yet, new questions and methodological challenges continue to emerge, including:

  • Accessibility. As platforms restrict data access—including shutting down APIs and prohibiting third-party data reuse—researchers face growing barriers to collecting the datasets needed for academic inquiry (Hu et al. 2025b).
  • Situatedness. Given the rapid turnover of online content and changes in platform algorithms, the availability and contextual meaning of digital traces shift constantly. This raises questions about how scholars can preserve contextual datasets for informed data interpretation (Koeser / LeBlanc 2024) (Hu et al. 2025a), particularly for studying niche, marginalized, or historically underinvestigated reading communities.
  • Data Quality. With the rise of trolling, coordinated backlash activities, and manipulated content (Hu et al. 2011), researchers must evaluate how representative, authentic, and reliable their datasets are.
  • Community Engagement. Existing research on user-generated book data often analyzes social-media data without contributors’ informed consent, though ethical research requires recognizing these individuals as identifiable participants who can experience real harm (e.g., cyber doxing, personal attacks). Yet widely varying disciplinary norms make it hard for researchers to define meaningful, ethical engagement (Hu et al. 2023), particularly when studying sensitive topics (e.g., user-led book bans) or marginalized readerships.

To address these challenges, this workshop brings together eight researchers to share their interdisciplinary approaches. Building on a previously convened panel that examined the opportunities and complexities of internet-based and book-related datasets, this workshop provides four training sessions with critical frameworks and actionable methods in support of research on book and reader communities using emerging materials.

DH-Workshop Leaders.

The eight workshop leaders, including both faculty members and PhD students, bring expertise spanning library science, literary studies, book history, cybersecurity and privacy, and computational social science. Their research explores critical methods and contextualized interpretation of born-digital datasets related to books, and reader/fan communities. Their publications address topics including the online circulation of poetry, social media discourse around challenged books, fanfiction, and the social reviewing of books across platforms. They also examine emerging challenges in digital culture, including misinformation, review manipulation, platform moderation, and intellectual freedom in social reviewing.

DH-Workshop Outcomes.

Drawing upon the instructors’ interdisciplinary expertise across platforms, languages, and culturally diverse online communities, the workshop will enable participants to (1) Critically conceptualize online books and reader datasets within their sociotechnical contexts; (2) Assess legal, ethical, community-based, and cybersecurity considerations involved in researching online datasets; (3) Develop research plans that support responsible preservation, curation, and critical analysis of these datasets; (4) Collaboratively design DH projects that integrate humanistic interpretation with critical data studies, emphasizing data accessibility and community engagement.

DH-Target Audience and Prerequisites.

This workshop is intended for DH researchers interested in online communities, contemporary book history, fan studies, and library science. It is especially relevant for scholars working with born-digital cultural materials, and/or conducting projects requiring ethical, contextualized, or interdisciplinary approaches. No prior technical background is required. Experience with Python may be helpful, but all programming-related content will be accessible to participants from non-technical backgrounds. The workshop is designed to support attendees with a broad range of methodological and technical experience.

DH-Workshop Agenda.

  • Session 0. Welcome & Orientation (15 minutes)
  • Session 1. Reviewing Authorship and Readership in the New Contexts (30 minutes): Social media reshapes the roles of authors, readers, and their dynamics. This session will guide participants to (1) discuss and reflect on authorial canonicity in the social-media age; (2) explore the affordances and limitations of platformed reader agency; and (3) use pre-existing datasets to envision research avenues from data collectives such as Post45 (Post45 n.d.).
  • Session 2. Managing Ethical and Practical Challenges: Fanfiction (45 minutes): Ethical challenges are central to research involving user‑generated data, as data collection, reuse, and analysis often intersect with issues of consent, authorship, and differing legal or institutional norms. Fanfiction stands between non-commercial creativity and increasingly monetized content, creating tensions among authors, fan writers, and researchers, which have been intensified with extensive data harvesting for AI.The session includes: (1) Overview of ethical challenges related to user-generated data. (2) Fanfiction Ethics: We will examine a fanfiction-related classroom assignment that sparked backlash (waldorph [Fanlore contributors] 2015), the Harry Potter fandom’s rejection of J.K. Rowling, and AO3’s evolving data policies. One instructor will share their process for assembling the Ficsim fanfiction dataset (Johnson et al. 2025), discussing informed consent, IRB approval, author privacy, and dataset licensing. (3) Exercise: Participants will collaboratively categorize a collection of fanfiction tags, contributing their knowledge of particular source media.
  • Coffee Break (15 minutes)
  • Session 3. Navigating Data Quality Challenges associated with Social Media Data (30 minutes): In recent years, incentivized, manipulated, or AI-generated content has become increasingly common on social media, often compromising the quality of cultural datasets collected from social media through data scraping or platform-provided APIs. In this session, participants will practice annotating and analyzing online book reviews that contain such content. Participants will learn how to identify “suspicious” content manually and scale up such detection work using dictionary-based computational methods, so as to improve the usability and reliability of social media datasets for scholarly research.
  • Session 4. Analyzing Large-Scale Online Bibliographic Metadata for Book History (30 minutes): Bibliographic metadata enables new forms of “distant reading” across large library collections, offering insights into “the great unread” and broader views of book history. This session will introduce methods for extracting information from library catalogs in the MARC (machine-readable cataloging) format—including descriptive metadata (e.g., title, language, and publication year) and subject information (e.g., subject heading and classification number)—to analyze book history and trace scholarly trajectories. Participants will be provided with an interactive Jupyter notebook containing prewritten Python code for hands-on activities, including parsing MARC records, cleaning data with regular expressions, and conducting data analysis and visualization.
  • Session 5. Reflection, Closing & Distribution of Workshop Materials (15 minutes)
References
  1. Hu, N. / Bose, I. / Gao, Y. / Liu, L. (2011): “Manipulation in digital word-of-mouth: A reality check for book reviews”, in: Decision Support Systems 50, 3: 627–635. DOI: 10.1016/j.dss.2010.08.013.
  2. Hu, Y. / Layne-Worthey, G. / Martaus, A. / Downie, J. S. / Diesner, J. (2023): “Research with User-Generated Book Review Data: Legal and Ethical Pitfalls and Contextualized Mitigations”, in: Sserwanga, I. / Goulding, A. / Moulaison-Sandy, H. / Du, J. T. / Soares, A. L. / Hessami, V. / Frank, R. D. (eds.): Information for a Better World: Normality, Virtuality, Physicality, Inclusivity. Cham: Springer Nature Switzerland 163–186. DOI: 10.1007/978-3-031-28035-1_13.
  3. Hu, Y. / Underwood, T. / Layne-Worthey, G. / Downie, J. S. (2025a): “Comparative analysis of classics book review data created by users across Douban and Goodreads”, in: Digital Scholarship in the Humanities 41, Supplement 1: i89–i106. DOI: 10.1093/llc/fqaf084.
  4. Hu, Y. / Diesner, J. / Underwood, T. / LeBlanc, Z. / Layne-Worthey, G. / Downie, J. S. (2025b): “Who decides what is read on Goodreads? Uncovering sponsorship and its implications for scholarly research”, in: Big Data & Society 12, 3: 1–17. DOI: 10.1177/20539517251359229.
  5. Johnson, N. / Bertsch, A. / Deal, M.-E. / Strubell, E. (2025): “FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction”, in: Christodoulopoulos, C. / Chakraborty, T. / Rose, C. / Peng, V. (eds.): Findings of the Association for Computational Linguistics: EMNLP 2025. Association for Computational Linguistics 25228–25246. DOI: 10.18653/v1/2025.findings-emnlp.1375.
  6. Koeser, R. S. / LeBlanc, Z. (2024): “Missing Data, Speculative Reading”, in: Journal of Cultural Analytics 9, 2. DOI: 10.22148/001c.116926.
  7. Murray, S. (2018): The Digital Literary Sphere: Reading, Writing, and Selling Books in the Internet Era. Baltimore: Johns Hopkins University Press.
  8. Murray, S. (2024): “Picking Your Professor: Bridging Scholarly and Popular Bookish Publics in the Digital Age”, in: Poetics Today 45, 4: 587–614. DOI: 10.1215/03335372-11393858.
  9. Post45 (n.d.): Post45 – American Literature and Culture since 1945. Website/database https://post45.org/ [01.05.2026].
  10. waldorph (Fanlore contributors) (2015): “So Your Fic is Required Reading: Hahahanope”, in: Fanlore.org https://fanlore.org/wiki/So_Your_Fic_is_Required_Reading:_Hahahanope [01.05.2026].
  11. Walsh, M. / Antoniak, M. (2021): “The Goodreads ‘Classics’: A Computational Study of Readers, Amazon, and Crowdsourced Amateur Criticism”, in: Journal of Cultural Analytics. DOI: 10.22148/001c.22221.
  12. Yu, Z. / Pianzola, F. (2024): “Across the Pages: A Comparative Study of Reader Response to Web Novels in Chinese and English on Qidian and WebNovel”, in: Proceedings of the Computational Humanities Research Conference 2024: 322–333 https://ceur-ws.org/Vol-3834/paper63.pdf [01.05.2026].