DH 2026

Daejeon, July 27–31

Thu, July 3015:20–16:20S112204-205
Short Paper

From Detection to Interpretation: Stylometry and AI Writing in Digital Humanities

Ronit Das
Indian Institute of Technology Guwahati, India · d.ronit@alumni.iitg.ac.in
Debapriya Basu
Indian Institute of Technology Guwahati, India · debapriya.basu@iitg.ac.in

Introduction

In November 2022, OpenAI launched ChatGPT, marking a significant shift in public-facing AI deployment (OpenAI 2022; Yu 2024). Since its release, ChatGPT has attracted sustained academic attention. Media and academic discourse have focused on its mechanisms, limitations, and educational implications (Shearing / McCallum 2023; Schulten 2023). This study, however, examines how ChatGPT reshapes core ideas of authorship and individual voice in academic writing.

Academic writing, in the humanities, has traditionally foregrounded personal expression and interpretive engagement. As LLMs like ChatGPT increasingly produce human-like prose, it becomes necessary to interrogate their implications for authorship and voice (OpenAI 2023; Dergaa et al. 2023). While their integration into higher education is often framed as supporting personalised learning and interactive pedagogy (Jebaselvi et al. 2024), concerns persist regarding ethical use, especially the risk of students submitting AI-generated work as their own (Lukianenko et al. 2024).

To address this, institutions have turned to AI-detection tools, yet recent studies question their reliability. (Elkhatat et al. 2023) document wide variation in detection accuracy, highlighting the difficulty of distinguishing AI-generated and human-written text. Similarly, (Ji et al. 2025) show that existing detectors lack explanatory transparency, while (Sadasivan et al. 2025) warn that high false-positive rates risk unjustly accusing students of AI plagiarism. These systems often misclassify human writing and offer little insight into why a text is flagged, thereby reinforcing opaque and often normative standards of “good writing.” Within Digital Humanities (DH), scholarship has primarily responded by either refining detection accuracy through quantitative methods (Kestemont 2014; Lukin et al. 2023) or by providing qualitative accounts of stylistic differences (Ringler 2022). Few DH studies meaningfully integrate these approaches or operationalise them within accessible, interpretive infrastructures.

Detection to Interpretation

Addressing this gap, this project proposes a mixed-methods DH framework that reorients scholarly engagement with AI-text analysis from detection toward interpretation. It combines quantitative stylometric analysis with qualitative narrative reading and implements this approach through a browser-based Digital Humanities tool, Voice Detect. Rather than classifying texts as human or machine, the tool avoids binary attribution and instead provides feature-level explanations that support interpretive engagement by scholars, students, and educators in examining how narrative voice is materially constructed and transformed across human and AI-generated prose.

Corpus and Research Questions

The study is based on a small, culturally specific corpus of ten paired argumentative essays written by Indian UPSC (Union Public Service Commission) aspirants, each approximately 900-950 words in length. The small corpus reflects both methodological design and practical constraints. Access to authentic UPSC essays proved difficult, as many aspirants were unwilling to share their work, often treating it as the product of intensive personal labour and as a form of intellectual property within a highly competitive examination context. This reluctance is itself indicative of the strong sense of ownership and authorial investment attached to such writing. Given the study’s interpretive focus, this limited corpus enables close, feature-level comparison rather than statistical generalisation. These essays were manually transcribed to preserve syntactic irregularities, stylistic and rhetorical texture. Parallel essays were generated using GPT-4o through controlled prompts that mirrored the original argumentative tasks and specified the aspirants’ educational backgrounds and rhetorical conditions, limiting variation across outputs. For instance, prompts took the form: “Consider yourself a UPSC aspirant with a background in [X discipline]. Write a (900-950 word) essay on ‘[X Topic]’, maintaining a formal, balanced argumentative style.” The UPSC genre, which focuses on ethics, public reasoning, governance, and national issues, constitutes a high-stakes civic writing context in which narrative voice carries social, cultural, and evaluative weight.

As one of India’s most competitive examinations, the UPSC has attracted sociological attention, with aspirants often investing years of intellectual and emotional labour. Despite its significance, this genre remains absent mainly from DH and LLM scholarship, which continues to privilege Western academic corpora.

By pairing localised human narratives with AI-generated counterparts, the project asks: How do LLMs engage with culturally embedded forms of argumentative writing? Which narrative features are preserved, regularised, or erased in the process of generation? These questions position AI not merely as a text-producing system, but as an active participant in shaping norms of academic expression.

Methodology

Methodologically, the analysis employs Voice Detect, a client-side, browser-based DH tool developed specifically for this project. Stylometry is the quantitative analysis of linguistic style for authorship attribution and identification of stylistic patterns (Eder et al. 2016). The tool allows users to upload paired texts and generates feature-level comparisons through an interactive interface. The analytical features are implemented through transparent, rule-based procedures applied consistently across texts, enabling comparative interpretation rather than precise measurement. The extracted features were reviewed to ensure consistency and interpretability.

It computes a curated set of lexical, syntactic, and discourse-level features, including lexical diversity, hedging, pronoun use, sentence structure, and readability indices. All processing occurs locally in the browser, ensuring transparency, reproducibility, and pedagogical accessibility without reliance on external APIs or black-box models. Importantly, these features are not treated as diagnostic indicators of authorship. Instead, they function as extensible interpretive annotations that illuminate how narrative identity and rhetorical stance are constructed differently by human writers and LLMs. Across the examined essay pairs, the analysis reveals recurring stylistic contrasts across the paired corpus. AI-generated essays tend to exhibit higher lexical evenness, increased hedging, syntactic regularity, reduced use of personal pronouns, and higher readability scores. Together, these features produce a polished, coherent, and globally normative academic voice. In contrast, human-authored essays display longer and more varied sentence structures, richer rhetorical ornamentation, broader conjunctional diversity, and greater irregularity, resulting in narratives that are more textured, situated, and culturally grounded.

Findings

These findings suggest that LLMs reproduce conservative models of “good writing” that prioritise clarity, neutrality, and balance, often at the cost of rhetorical specificity and situated voice. From a DH perspective, this raises critical concerns. Pedagogically, students may internalise depersonalised, standardised prose as the ideal academic form. At the level of narrative identity, human rhetorical textures risk being overshadowed by machine-preferred styles. In terms of knowledge production, AI-generated writing often prioritises neutrality over lived experience. From the standpoint of epistemic justice, voices from the Global South, such as those represented in UPSC essays, risk being assimilated into homogenised narrative templates shaped by Anglophone academic norms.

Contribution

By foregrounding small data, interpretive annotation, and tool-based transparency, this project contributes to DH debates on engagement at the human-AI interface in three key ways. First, it adds to existing scholarship on how stylometry can function as an interpretive practice rather than a primarily classificatory one. Second, it offers a reusable, browser-based analytical interface that makes computational analysis legible to humanities scholarswhile grounded in the UPSC essay genre, the framework is adaptable to other cultural contexts and writing genres, enabling broader comparative DH research on AI-mediated authorship. Finally, it reframes engagement with AI as a critical, reflective DH practice, one that examines not only how machines write, but how their writing reshapes evolving understandings of authorship, creativity, and voice within digital culture.

References
  1. Dergaa, Ismail / Chamari, Karim / Zmijewski, Piotr / Saad, Helmi Ben (2023): “From human writing to artificial intelligence generated text: examining the prospects and potential threats of ChatGPT in academic writing”, in: Biology of Sport 40, 2: 615–622.
  2. Eder, Maciej / Rybicki, Jan / Kestemont, Mike (2016): “Stylometry with R: A Package for Computational Text Analysis”, in: The R Journal 8, 1: 117–121.
  3. Elkhatat, Ahmed M. / Elsaid, Khaled / Almeer, Saeed (2023): “Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text”, in: International Journal for Educational Integrity 19: 1–16.
  4. Jebaselvi, C. Alice Evangaline / Mohanraj, K. / Anitha, T. (2024): “The rise of AI in English Language and Literature”, in: Shanlax International Journal of English 12, 2: 53–58.
  5. Ji, Jiazhou / Li, Ruizhe / Li, Shujun / Guo, Jie / Qiu, Weidong / Huang, Zheng / Chen, Chiyu / Jiang, Xiaoyu / Lu, Xinru (2025): “Detecting machine generated texts: Not just “AI vs Human” and explainability is complicated”, in: arXiv preprint arXiv: 2406.18259 <https://doi.org/10.48550/arXiv.2406.18259> [07.05.2026].
  6. Kestemont, Mike (2014): “Macroanalysis. Digital Methods and Literary History. Matthew L. Jockers”, in: Literary and Linguistic Computing 29, 2: 274–276.
  7. Lukianenko, Valentyna Volodymyrivna / Shastko, Iryna Myronivna / Korbut, Oksana Hryhorivna (2024): “Evaluating AI Detection Tools for Academic Integrity in Higher Education”, in: Naukovi innovatsii ta peredovi tekhnolohii 5, 33: 970–978. DOI: 10.52058/2786-5274-2024-5(33)-970-978.
  8. Lukin, Eugenia / Roberts, James Cooper / Berdik, David / Mugar, Eliana / Juola, Patrick (2023): “Adjectives and adverbs as stylometric analysis parameters”, in: International Journal of Digital Humanities 5: 233–245.
  9. OpenAI (2022): “Introducing ChatGPT”, in: OpenAI, 30 November 2022 <https://openai.com/index/chatgpt/> [07.05.2026].
  10. OpenAI (2023): “GPT-4 Technical Report”, in: arXiv preprint arXiv: 2303.08774 <https://arxiv.org/abs/2303.08774> [07.05.2026].
  11. Ringler, Hannah (2022): “’We can’t read it all’: Theorizing a hermeneutics for large-scale data in the humanities with a case study in stylometry”, in: Digital Scholarship in the Humanities 37, 4: 1157–1171.
  12. Sadasivan, Vinu Sankar / Kumar, Aounon / Balasubramanian, Sriram / Wang, Wenxiao / Feizi, Soheil (2025): “Can AI-Generated Text be Reliably Detected?”, in: arXiv preprint arXiv: 2303.11156 <https://doi.org/10.48550/arXiv.2303.11156> [07.05.2026].
  13. Schulten, Katherine (2023): “Lesson Plan: Teaching and Learning in the Era of ChatGPT”, in: The New York Times, 24 January 2023 <https://www.nytimes.com/2023/01/24/learning/lesson-plans/lesson-plan-teaching-and-learning-in-the-era-of-chatgpt.html> [07.05.2026].
  14. Shearing, Hazel / McCallum, Shiona (2023): “ChatGPT: Can students pass using AI tools at university?”, in: BBC News, 9 May 2023 <https://www.bbc.com/news/education-65316283> [07.05.2026].
  15. Yu, Hao (2024): “The application and challenges of ChatGPT in educational transformation: New demands for teachers’ roles”, in: Heliyon 10, 2: e24289.