DH 2026

Daejeon, July 27–31

Poster

Do Readers Accept AI as an Author? A Preliminary Analysis of Reddit Discussions on AI-Generated Book Content

Owen Kleinmaier
Luddy School of Informatics, Computing and Engineering, Indiana University Bloomington, United States of America · owklein@iu.edu
Yuerong Hu
Luddy School of Informatics, Computing and Engineering, Indiana University Bloomington, United States of America · yuerhu@iu.edu

Increasing AI-generated content in books has sparked concern and pushback, as many believe these books harm the interest of human authors and erode trust in the literary marketplace (Goodfellow 2024; Sato 2023). Such concerns have prompted real-world action in the publishing industry, such as certifications for “human-author-only” books (Authors Guild 2025; Loffhagen 2025). On social reviewing platforms, some readers have called for boycotts of books with AI-generated covers—viewed as unethical derivatives of human labor—while others argue that such backlash unfairly harms authors, particularly when production decisions are made by publishers (Goodreads 2023). These tensions and debates call for critical investigation into the expansion of generative book content, and the perceived sociotechnical challenges it has introduced.

To this end, this research explores how books with AI-generated content have been discussed and contested online, through a case study based on data collected from Reddit, a social platform for community-based and topic-driven discussion with abundant textual data for our research questions (Proferes et al. 2021). We combined computational text analysis with qualitative investigation to analyze the data.

Figure 1 shows our workflow, from data collection to mixed-method analysis. We collected 775,640 comments in English via Reddit's public API using 16 book-related search terms. Figure 2 presents the search terms used, and their corresponding numbers of posts collected. These terms were selected on the basis of our prior research on online book reviews (Hu et al. 2023; Hu et al. 2025a; Hu et al. 2025b). The list was not exhaustive; instead, it was designed to capture a representative range of book-related discussions for exploratory analysis. In the next stage of the study, we will expand and refine the search strategy to include discussions on a broader range of platforms and topics.

With both LDA topic modeling (McCallum 2002) and dictionary-based keyword matching, we identify 386 posts about books with AI-generated content. These selected posts were tagged as “AI book posts” while the rest of the posts were tagged as “general book posts”. Next, we conducted VADER sentiment analysis (Hutto & Gilbert 2014) to compare the sentiment of AI book posts and general book posts, and conducted qualitative analysis of content (Wildemuth 2016) on the AI book posts.

Figure 1. Overall workflow for data collection and processing

Figure 2. Distribution of collected Reddit text data across 16 search terms

Quantitative analysis revealed that AI book posts (n=386) showed more negative sentiment than general book posts (n=277,323), with mean VADER scores of 0.115 and 0.180, respectively (t = -2.57, p = 0.01). While this difference is statistically significant, the effect size is small (Cohen's d = -0.131). To examine whether this result was influenced by the substantial imbalance in dataset sizes, we computed sentiment scores for 1,000 samples of 386 randomly selected general book posts. In these tests, AI book posts still showed more negative sentiment, but only 28.7% of tests reached statistical significance. This suggests that the original result may be partially influenced by the imbalanced sizes of data. For meaningful comparisons, we plan to enlarge the existing AI book posts datasets by collecting relevant posts from other platforms. We also plan to conduct sentiment analysis with word embeddings as VADER’s dictionary-based approach may miss nuanced affective expressions such as sarcasm.

We complemented our quantitative analysis with qualitative analysis and identified four major clusters of topics: (1) human value, where readers raise ethical questions about what counts as human-made art and creative works; (2) labor concerns, reflecting concerns about AI replacing creative professionals such as book authors, editors, and illustrators; (3) authenticity, with AI writing viewed as imitation rather than genuine creativity; and (4) calls for disclosure, with many insisting AI-generated books must be clearly labeled to enable informed and ethical consumption. Figure 3 presents two exemplary quotes for each cluster. Across these clusters, most of the posts challenge the legitimacy of AI authorship and express ethical concerns surrounding AI-generated content.

It is important to note that even negative critiques of AI-generated book content tend to be more evaluative and discussive, rather than completely negative or hostile. Additionally, a few readers expressed openness or neutrality to AI books. For example, one user posted that “Using AI to assist in writing is very much a ‘you get out of it what you put’ sort of relationship”, and another user stated, “I am not against an author having an AI editor”.

Figure 3. Representative quotes across four major clusters of topics in AI book posts

Our mixed-method analysis of Reddit data reveals sustained concerns in discussions about AI-generated book content, with a few users viewing AI as a potentially useful tool. To deepen and enrich this preliminary investigation, we plan to collect posts on generative book content across platforms and user bases. In particular, recognizing the cultural and linguistic dependencies of book discourse (Hu et al. 2025a), we aim to include book discourse in languages other than English and generated by geo-culturally diverse readerships.

References
  1. Authors Guild (2025): “Authors Guild launches ‘Human Authored’ certification”, in: Authors Guild https://authorsguild.org/human-authored/ [01.05.2026].
  2. Bender, Stuart (2025): “Generative-AI, the media industries, and the disappearance of human creative labour”, in: Media Practice and Education 26, 2: 200–217. DOI: 10.1080/25741136.2024.2355597.
  3. Goodfellow, P. (2024): “The Distributed Authorship of Art in the Age of AI”, in: Arts 13, 5: 149. DOI: 10.3390/arts13050149.
  4. Goodreads (2023): “Fractal Noise”, book page on Goodreads. Social reviewing platform https://www.goodreads.com/book/show/62711641-fractal-noise [01.05.2026].
  5. Hu, Y. / Layne-Worthey, G. / Martaus, A. / Downie, J. S. / Diesner, J. (2023): “Research with User-Generated Book Review Data: Legal and Ethical Pitfalls and Contextualized Mitigations”, in: Sserwanga, I. / Goulding, A. / Moulaison-Sandy, H. / Du, J. T. / Soares, A. L. / Hessami, V. / Frank, R. D. (eds.): Information for a Better World: Normality, Virtuality, Physicality, Inclusivity. Cham: Springer Nature Switzerland 163–186. DOI: 10.1007/978-3-031-28035-1_13.
  6. Hu, Y. / Underwood, T. / Layne-Worthey, G. / Downie, J. S. (2025a): “Comparative analysis of classics book review data created by users across Douban and Goodreads”, in: Digital Scholarship in the Humanities. Advance online publication: fqaf084. DOI: 10.1093/llc/fqaf084.
  7. Hu, Y. / Diesner, J. / Underwood, T. / LeBlanc, Z. / Layne-Worthey, G. / Downie, J. S. (2025b): “Who decides what is read on Goodreads? Uncovering sponsorship and its implications for scholarly research”, in: Big Data & Society 12, 3: 1–17. DOI: 10.1177/20539517251359229.
  8. Hutto, C. J. / Gilbert, E. (2014): “VADER: A parsimonious rule-based model for sentiment analysis of social media text”, in: Proceedings of the International AAAI Conference on Web and Social Media 8, 1: 216–225. DOI: 10.1609/icwsm.v8i1.14550.
  9. Loffhagen, E. (2025): “Certified organic and AI-free: New stamp for human-written books launches”, in: The Guardian, 15 October 2025 https://www.theguardian.com/books/2025/oct/15/books-by-people-for-people-publishers-launch-certification-human-written-ai [01.05.2026].
  10. McCallum, Andrew Kachites (2002): MALLET: A Machine Learning for Language Toolkit. Software http://mallet.cs.umass.edu [01.05.2026].
  11. Proferes, N. / Jones, N. / Gilbert, S. / Fiesler, C. / Zimmer, M. (2021): “Studying Reddit: A systematic overview of disciplines, approaches, methods, and ethics”, in: Social Media + Society 7, 2: 1–14. DOI: 10.1177/20563051211019004.
  12. Sato, M. (2023): “AI-generated fiction is flooding literary magazines—but not fooling anyone”, in: The Verge, 25 February 2023 https://www.theverge.com/2023/2/25/23613752/ai-generated-short-stories-literary-magazines-clarkesworld-science-fiction [01.05.2026].
  13. Wildemuth, B. M. (ed.) (2016): Applications of Social Research Methods to Questions in Information and Library Science. 2nd ed. Santa Barbara, CA: Libraries Unlimited.