Daejeon, July 27–31
This presentation reports on an experiment investigating the perception of fan fiction by non-professional readers, both published on the Japanese platform Shōsetsuka ni Narō (https://syosetu.com) and generated by a custom fine-tuned Large Language Model (LLM) fine-tuned for short fiction generation within limited computational resources as well as a commercial general use LLM (ChatGPT-5.1). In this research, we designed a compact survey with respondents majoring in subjects other than Literature at Kyoto Sangyo University, Japan.
Our study challenges the bias that AI is inferior to a human, either in practical tasks or in creativity. The AI, generating “slop” content such as short videos, occupied a significant part of SNS users content consumption. Although AI is not an outstanding author, it proved that it can entertain despite the bias existing against it. The results of our survey demonstrate that in terms of what readers value in texts, they can robustly distinguish between human and AI neither in text attribution nor in literary features.
Although there is increasing interest in storytelling generation with LLMs (El-Shangiti et al. 2024; Bangura et al. 2023; Beguš 2024; Yuan et al. 2022; Elias et al. 2025; Rosa et al. 2021), full-scale studies on the perception of such texts by readers are still limited. Research based on questionnaires or qualitative methodologies focus on LLM outputs in essay format (Doru et al. 2025; Fiedler / Döpke, 2025). Both close-reading studies of LLM-created (Beguš 2024) fiction and questionnaire studies based on reader response methodologies remain limited (Raffloer and Green, 2025; Nakano et. al. 2025). Current DH also tends to look at LLMs as a tool of distant writing or expansion of frequency-based DH methodologies rather than on the object of content which is consumed by readers.
After filtering to fit the short story texts with generation prompt within 2,048 tokens model’s limitations, training data for fine-tuning of LLM counted 110,203 short stories.
Training and fine-tuning the model required significant computational resources. Given ongoing concerns within the Digital Humanities community about ensuring accessibility of tools, and considering our own computational limitations (specifically, fine-tuning is planned on a GPU with 16 GB of VRAM), we selected the bilingual-gpt-neox-4b model — a bilingual English–Japanese transformer LLM — for fine-tuning (Sawada et al. 2024). This choice reflects the limited availability of lightweight, open-access LLMs for Japanese. To optimize fine-tuning under resource constraints, we used parameter-efficient fine-tuning (PEFT) with Low-Rank Matrix Adaptations (LoRA) (Hu et al. 2021).
Prompts instructed the model to generate short stories by genre, keyword tags, and title, extracted from the metadata.
The respondents (N=35) received a questionnaire consisted of two sets of questions for two stories, totally of the following structure (Table 1).
| Question Group | Number |
| General Questions (age, reading habits, AI usage habits) | 14 |
| Overall Impression (correspondence of title/keywords with the content, attitude to the text) | 4 x 2 |
| Language and Style (readability, literary quality, presence of individual author’s style) | 7 x 2 |
| Plot (originality, empathy to the story action, appropriateness or offensiveness) | 10 x 2 |
| Characters (empathy to characters, completeness of character’s image) | 6 x 2 |
| Structure (visual organization of text) | 3 x 2 |
| Dialogue (natural development, coherence with the plot, logic) | 3 x 2 |
| After Reading (general expression, desire to continue reading) | 4 x 2 |
Table 1. Questionnaire structure.
In the main section, each participant reads two short stories in one of three possible compositions: 1) both stories are written by human, 2) both stories are generated by AI (LLM or by the commercial LLM (ChatGPT-5.1), matched by title, genre, and keywords), 3) one story is by human and one story is by AI. The set contains 24 pairs of short stories. The short stories of the same title, genre, and keywords are displayed to different readers both in AI and human authorship, which is disclosed at the end of the questionnaire. Each pair of texts is evaluated by two respondents from the base and control groups in the Likert scale from 1 to 5.
Ethical considerations include informed consent, anonymity and confidentiality of data, and the absence of any interventions posing physical, psychological, or legal risks. Participants may stop the questionnaire at any time before or during the process.
In terms of the respondents’ ability to guess the writer of the short story, the ChatGPT 5.1 commercial model demonstrated better mimicry of human texts. In the case of the custom model, respondents were never mistaken in finding AI-written text; however, predominantly they were more suspicious of human texts being written by AI. The corpus composition contributed to this suspicion as it mixed texts by high-ranked and low-ranked authors, so meeting a “weak” story, respondents prescribed it to AI. This respondent’s behavior highlights their existing bias against AI with a practical inability to robustly distinguish between human and AI stories.
| Configuration | Precision | Accuracy | F1 Score |
| Human vs. Fine-Tuned Model | 1 | 0.55 | 0.71 |
| Human vs. ChatGPT 5.1 | 0.52 | 0.26 | 0.35 |
| Human vs. both models | 0.52 | 0.45 | 0.48 |
Table 2. Precision, accuracy, and F1 score of respondents’ guessing the authorship.
A Repeated Measures ANOVA (Analysis of Variance) of the responses revealed that there is no statistically significant difference between human and AI texts in their textual features, regardless of the stated authorship (Table 3). Consequently, in practical terms, respondent responses do not provide sufficient information to distinguish between the three categories of texts.
| Human vs Fine-Tuned Model | Human vs ChatGPT | |||||||
| F Value | Num DF | Den DF | Pr > F | F Value | Num DF | Den DF | Pr > F | |
| Authorship | 0.4385 | 1 | 4 | 0.5441 | 1.2985 | 1 | 6 | 0.2979 |
| Question Group Category | 3.5591 | 6 | 24 | 0.0115 | 11.9046 | 6 | 36 | 0.001 |
| Authorship vs. Question Group Category | 1.3228 | 6 | 24 | 0.2854 | 1.1911 | 6 | 36 | 0.3331 |
Table 3. Repeated Measure ANOVA results.
A post hoc comparison between question categories and authorship did not reveal any statistically significant difference between the two categories of texts. The fine-tuned model demonstrated a diminished ability to generate high-quality dialogues and coherent, captivating plots compared to human-written texts. However, surprisingly, respondents did not perceive a significant difference in language quality, specifically in terms of readability and the absence of author’s style. Notably, respondents observed that the models’ plots were slightly superior (Table 4).
| Human vs. ChatGPT | |||||
| question_category | Human_mean | AI_ChatGPT_mean | mean_difference | paired_t_p_corrected | wilcoxon_p_corrected |
| after_reading | 2.214285714 | 2.071428571 | 0.1428571429 | 0.8378155686 | 1 |
| characters | 2.804945055 | 3.142857143 | -0.3379120879 | 0.6254570642 | 1 |
| dialogue | 3.5 | 3.880952381 | -0.380952381 | 0.6254570642 | 0.546875 |
| language_style | 2.671428571 | 2.774025974 | -0.1025974026 | 0.8378155686 | 1 |
| overall_impression | 2.814285714 | 3.428571429 | -0.6142857143 | 0.6254570642 | 0.4375 |
| plot | 2.820903361 | 2.839285714 | -0.01838235294 | 0.8378155686 | 1 |
| structure | 3.233333333 | 3.404761905 | -0.1714285714 | 0.6254570642 | 0.6927083333 |
| Human vs. Fine-Tuned Model | |||||
| question_category | Human_mean | Fine-tuned_mean | mean_difference | paired_t_p_corrected | wilcoxon_p_corrected |
| after_reading | 1.9 | 2.1 | -0.2 | 0.826800794 | 1 |
| characters | 2.85 | 2.416666667 | 0.4333333333 | 0.6353527539 | 0.875 |
| dialogue | 3.166666667 | 2.766666667 | 0.4 | 0.6353527539 | 0.875 |
| language_style | 2.861010101 | 2.898181818 | -0.03717171717 | 0.8477445763 | 1 |
| overall_impression | 3.088888889 | 2.633333333 | 0.4555555556 | 0.6353527539 | 0.875 |
| plot | 2.9875 | 3.085833333 | -0.09833333333 | 0.826800794 | 1 |
| structure | 3.033333333 | 2.833333333 | 0.2 | 0.826800794 | 1 |
Table 4. Ad hoc analysis of question categories.
The texts by ChatGPT received higher scores in every group except for after reading impressions. The commercial model created the text close to the prompt following keywords, genre, and title. Although its texts received higher scores in after-reading impressions, they were closer to one-time-reading (as well as human-written texts).
Both models produced results that reflect broader transformations in the mass culture of industrial and post-industrial society. Within platform-based leisure reading, fiction increasingly functions as consumable cultural content, and AI-generated texts appear capable of fulfilling this role no less effectively than human authors by producing coherent plots, vivid dialogues, and narratives aligned with reader expectations.
At the same time, the responses suggest that readers may still perceive an affective quality in texts that cannot be fully reduced to technical fluency alone. The commercially developed OpenAI model generated smoother and more polished outputs yet showed slightly weaker after-reading impressions. Conversely, the comparatively rough fine-tuned fiction model occasionally produced stronger after-reading effects despite its imperfections. These findings raise the possibility that future fiction-oriented models may be trained not only to imitate narrative structures, but also to evoke forms of emotional engagement traditionally associated with human writing.