DH-Workshop Topic and Introduction.
In the last two decades, unprecedented volumes of data about books and reader communities have been available online, such as user-added genre tags (Walsh / Antoniak 2021), “BookTok” videos (Murray 2024), fanfiction (Johnson et al. 2025), and web novels (Yu / Pianzola 2024). These internet datasets have been increasingly used for digital humanities (DH) research on reading, literature reception, and contemporary book culture (Murray 2018). Yet, new questions and methodological challenges continue to emerge, including:
- Accessibility. As platforms restrict data access—including shutting down APIs and prohibiting third-party data reuse—researchers face growing barriers to collecting the datasets needed for academic inquiry (Hu et al. 2025b).
- Situatedness. Given the rapid turnover of online content and changes in platform algorithms, the availability and contextual meaning of digital traces shift constantly. This raises questions about how scholars can preserve contextual datasets for informed data interpretation (Koeser / LeBlanc 2024) (Hu et al. 2025a), particularly for studying niche, marginalized, or historically underinvestigated reading communities.
- Data Quality. With the rise of trolling, coordinated backlash activities, and manipulated content (Hu et al. 2011), researchers must evaluate how representative, authentic, and reliable their datasets are.
- Community Engagement. Existing research on user-generated book data often analyzes social-media data without contributors’ informed consent, though ethical research requires recognizing these individuals as identifiable participants who can experience real harm (e.g., cyber doxing, personal attacks). Yet widely varying disciplinary norms make it hard for researchers to define meaningful, ethical engagement (Hu et al. 2023), particularly when studying sensitive topics (e.g., user-led book bans) or marginalized readerships.
To address these challenges, this workshop brings together eight researchers to share their interdisciplinary approaches. Building on a previously convened panel that examined the opportunities and complexities of internet-based and book-related datasets, this workshop provides four training sessions with critical frameworks and actionable methods in support of research on book and reader communities using emerging materials.
DH-Workshop Leaders.
The eight workshop leaders, including both faculty members and PhD students, bring expertise spanning library science, literary studies, book history, cybersecurity and privacy, and computational social science. Their research explores critical methods and contextualized interpretation of born-digital datasets related to books, and reader/fan communities. Their publications address topics including the online circulation of poetry, social media discourse around challenged books, fanfiction, and the social reviewing of books across platforms. They also examine emerging challenges in digital culture, including misinformation, review manipulation, platform moderation, and intellectual freedom in social reviewing.
DH-Workshop Outcomes.
Drawing upon the instructors’ interdisciplinary expertise across platforms, languages, and culturally diverse online communities, the workshop will enable participants to (1) Critically conceptualize online books and reader datasets within their sociotechnical contexts; (2) Assess legal, ethical, community-based, and cybersecurity considerations involved in researching online datasets; (3) Develop research plans that support responsible preservation, curation, and critical analysis of these datasets; (4) Collaboratively design DH projects that integrate humanistic interpretation with critical data studies, emphasizing data accessibility and community engagement.
DH-Target Audience and Prerequisites.
This workshop is intended for DH researchers interested in online communities, contemporary book history, fan studies, and library science. It is especially relevant for scholars working with born-digital cultural materials, and/or conducting projects requiring ethical, contextualized, or interdisciplinary approaches. No prior technical background is required. Experience with Python may be helpful, but all programming-related content will be accessible to participants from non-technical backgrounds. The workshop is designed to support attendees with a broad range of methodological and technical experience.
DH-Workshop Agenda.
- Session 0. Welcome & Orientation (15 minutes)
- Session 1. Reviewing Authorship and Readership in the New Contexts (30 minutes): Social media reshapes the roles of authors, readers, and their dynamics. This session will guide participants to (1) discuss and reflect on authorial canonicity in the social-media age; (2) explore the affordances and limitations of platformed reader agency; and (3) use pre-existing datasets to envision research avenues from data collectives such as Post45 (Post45 n.d.).
- Session 2. Managing Ethical and Practical Challenges: Fanfiction (45 minutes): Ethical challenges are central to research involving user‑generated data, as data collection, reuse, and analysis often intersect with issues of consent, authorship, and differing legal or institutional norms. Fanfiction stands between non-commercial creativity and increasingly monetized content, creating tensions among authors, fan writers, and researchers, which have been intensified with extensive data harvesting for AI.The session includes: (1) Overview of ethical challenges related to user-generated data. (2) Fanfiction Ethics: We will examine a fanfiction-related classroom assignment that sparked backlash (waldorph [Fanlore contributors] 2015), the Harry Potter fandom’s rejection of J.K. Rowling, and AO3’s evolving data policies. One instructor will share their process for assembling the Ficsim fanfiction dataset (Johnson et al. 2025), discussing informed consent, IRB approval, author privacy, and dataset licensing. (3) Exercise: Participants will collaboratively categorize a collection of fanfiction tags, contributing their knowledge of particular source media.
- Coffee Break (15 minutes)
- Session 3. Navigating Data Quality Challenges associated with Social Media Data (30 minutes): In recent years, incentivized, manipulated, or AI-generated content has become increasingly common on social media, often compromising the quality of cultural datasets collected from social media through data scraping or platform-provided APIs. In this session, participants will practice annotating and analyzing online book reviews that contain such content. Participants will learn how to identify “suspicious” content manually and scale up such detection work using dictionary-based computational methods, so as to improve the usability and reliability of social media datasets for scholarly research.
- Session 4. Analyzing Large-Scale Online Bibliographic Metadata for Book History (30 minutes): Bibliographic metadata enables new forms of “distant reading” across large library collections, offering insights into “the great unread” and broader views of book history. This session will introduce methods for extracting information from library catalogs in the MARC (machine-readable cataloging) format—including descriptive metadata (e.g., title, language, and publication year) and subject information (e.g., subject heading and classification number)—to analyze book history and trace scholarly trajectories. Participants will be provided with an interactive Jupyter notebook containing prewritten Python code for hands-on activities, including parsing MARC records, cleaning data with regular expressions, and conducting data analysis and visualization.
- Session 5. Reflection, Closing & Distribution of Workshop Materials (15 minutes)