DH 2026

Daejeon, July 27–31

Thu, July 3015:20–16:20S112204-205
Short Paper

Can LLMs imitate an author's style? Comparative stylometric study of ChatGPT, Gemini, and Claude.

Aleksandra Rykowska
Jagiellonian University in Kraków, Poland · aleksandra.rykowska@doctoral.uj.edu.pl

Since November 2022, when OpenAI released ChatGPT to the public, the linguistic world has been revolutionized by the omnipresence of large language models. Despite their popularity in linguistic research, there exists only a handful of studies that analyze the stylistics of automatically generated texts, especially in terms of literary language. A very prominent theme in AI research is the perception of automatically generated texts by humans. As proven, AI-generated poems are indistinguishable from poetry written by humans (Porter and Machery, 2024), and in terms of journalistic style, LLMs proved to generate content of high linguistic quality (Zhang et al., 2024). Agrahari et al. (2024) “underscore the need for better detection techniques to distinguish between human and AI-generated text effectively.”

On the other hand, Amirjalili et al. (2024) highlight that in the case of academic writing, wherein the generated text might “seem human-like at first glance, but lacks generic style.” Similarly, Bhandarkar et al. (2024) investigated pre-trained LLMs for author style emulation and found that the LLMs do not perform the task as smoothly as expected, and Przystalski et al. (2025) state that it is possible to distinguish human and AI-generated texts using stylometry even on small samples. O’Sullivan (2025) presented a similar case: a stylometric study of human- and automatically generated creative writing that was easily distinguishable from one another.

The most recent study by Mikros (2025) examines GPT-4o’s ability to imitate the styles of Hemingway and Shelley. The author states that, even though it may seem that GPT can fool stylometry, it fails to truly imitate the original authors at a deeper linguistic level. Given the robust research on the stylometry of human vs. automatically generated texts, which yields conflicting results, it is necessary to investigate the issue further. LLM’s ability to imitate the style of a specific author might additionally raise ethical and legal questions.

This study proposes to examine the ability of popular LLMs (GPT, Claude, Gemini) to imitate the authorial style of Leopold Tyrmand, a widely recognized author of prose novels set in the period of the Polish People’s Republic. While this writer is mostly recognized as the author of “The Man with White Eyes,” Tyrmand was prominent in journalism, exposing the absurdities of communism. He also kept a diary during the first months of 1954, which is the subject of this study. “Journal 1954” has been chosen for this study for several reasons: A diary or journal is a text that is in principle a collage of separate, shorter excerpts, making it a suitable material for the analysis of authorial style. While some discuss that it was created to imitate the style of a novel, and to create Tyrmand himself as a novel's protagonist (Niciński 2006), a recent study revealed that stylometrically, the journal is more similar to Tyrmand's journalistic writing rather than his novels (Rykowska 2023).

While it is easier to find Tyrmand's novels on the Internet, there is no full version of the journal that has been published in open access – this might pose a greater challenge for AI to imitate the text. The automatically generated diary entries were created using zero-shot prompting. Each model generated 20 journal entries, approximately 2000 words each, set in the first three months of 1954. The dates were randomly chosen.

Two stylometric tests were performed using the stylo framework (Eder et al. 2016). The first one was a typical Bootstrap consensus of 100–500 most frequent words using the cosine delta. The results were then visualized using Gephi (Bastian et al. 2009). For this test, each automatically generated text was saved in a separate file. 20 randomly chosen diary entries written by Tyrmand were also selected and put into separate files. The second test, using the rolling stylometry method, included a study in which the generated entries were placed within the original journal (in accordance with the chronology). Three tests using the SVM, NSC, and Delta methods (Eder 2016) were performed. The start of the automatically generated texts was marked as a milestone to facilitate the reading of each chart.

Apart from limiting the study to MFW, several analyses of the linguistic layer of the texts were conducted to complement the study. The texts were tokenized using the NLTK libraries for Python (Bird et al. 2009) on the sentence level, and sentence length variability was established. The texts were lemmatized using the spacy pl_core_news_sm pipeline (Honnibal et al. 2020). Moving TTR and vocabulary growth were calculated, and two readability tests were performed, using the Flesch Reading Ease and the Flesch-Kincaid Grade. All generated texts and the analysis results are shared in the GitHub repository at: https://github.dev/TyrmandStylo/DH2026.

The test indicates that the three models fail to imitate Leopold Tyrmand’s journalistic style. However, not all of the above-mentioned methods proved unsuccessful. While the NSC method does not indicate the automatically generated fragments of “Journal 1954” (see Fig. 1), the SVM method captures the stylistic change in almost every case (see Fig. 2).

Figure Figure 1. NSC classification of the "Journal 1954" with automatically generated by Gemini entries.
Figure Figure 1. NSC classification of the "Journal 1954" with automatically generated by Gemini entries. — Figure 2. SVM classification of the "Journal 1954" with automatically generated by GPTi entries.

Cluster analysis and Bootstrap consensus clearly show the stylistic differences not only between human and automatically generated texts, but also between different models (see Figs. 3 and 4).

Figure Figure 3. Cluster analysis of the generated texts and Tyrmand's diary entries.
Figure Figure 3. Cluster analysis of the generated texts and Tyrmand's diary entries. — Figure 4. Bootsrap consensus tree of the generated texts and Tyrmand's diary entries.

The study aims to analyze the ability of popular large language models to imitate an author's literary style. The results of various stylometric and linguistic studies show that it is still possible to automatically distinguish the styles of humans and LLMs, as well as the different LLMs used to generate the texts, which is an addition to the existing state of the art.

References
  1. Agrahari, Shifali / Bisht, Samridhi / Sanasam Ranbir Singh (2024): “Text Authorship Attribution : Stylometric Insights into Human and LLM-Generated Text”, in: CODS-COMAD.
  2. Amirjalili, Forough / Neysani, Masoud / Nikbakht, Ahmadreza (2024): “Exploring the boundaries of authorship: a comparative analysis of AI-generated text and human academic writing in English literature”, in: Frontiers in Education 9:1347421.
  3. Bastian, Mathieu / Heymann, Sebastien / Jacomy, Mathieu (2009): “Gephi: an open source software for exploring and manipulating networks”, in: International AAAI Conference on Weblogs and Social Media. San Jose, California. DOI: 10.13140/2.1.1341.1520.
  4. Bhandarkar, Avanti et al (2024): “Emulating Author Style: A Feasibility Study of Prompt-enabled Text Stylization with Off-theShelf LLMs”, in: Proceedings of the 1st Workshop of Personalization of Generative AI Systems (PERSONALIZE 2024). Association for Computational Linguistics, 76-82.
  5. Bird, Steven / Klein, Ewan / Loper. Edward (2009): Natural Language Processing with Python, O'Reilly Media Inc.
  6. Eder, Maciej / Rybicki, Jan / Kestemont, Mike (2016): “Stylometry with R: a package for computational text analysis”, in: The R Journal 8(1): 107–121.
  7. Eder, Maciej (2016): “Rolling stylometry”, in: Digital Scholarship in the Humanities, 31(3): 457–469.
  8. Honnibal, Matthew et al. (2020): spaCy : Industrial-strength Natural Language Processing in Python. DOI: 10.5281/zenodo.1212303
  9. Mikros, George (2025): “Beyond the surface: stylometric analysis of GPT-4o’s capacity for literary style imitation”, in: Digital Scholarship in the Humanities, 40: 587-600.
  10. Niciński, Konrad (2006): “Dwie wersje Dziennika 1954 Leopolda Tyrmanda. Wokół problem tożsamości tekstu”, in: Pamiętnik Literacki 4: 71–94.
  11. O’Sullivan, James (2025): “Stylometric comparisons of human versus AI-generated creative writing”, in: Humanities & Social Sciences Communications 12: 1708.
  12. Porter, Brian / Machery, Edouard (2024): “AI-Generated Poetry is Indistinguishable from Human-Written Poetry and Is Rated More Favorably”. Scientific Reports 14: 26133.
  13. Przystalski, K. et al. (2025): “Stylometry recognizes humand and LLM-generated texts in short samples”. Preprint submitted to Expert Systems with Applications.
  14. Rykowska, Aleksandra (2023): “Językowy obraz komunizmu w Dzienniku 1954 Leopolda Tyrmanda. Interpretacja wspomagana metodami i narzędziami lingwistyki komputerowej”, in: Annales Universitatis Paedagogicae Cracoviensis. Studia de Cultura 15(2).
  15. Zhang, Tianyi et al . (2024): “Benchmarking Large Language Models for News Summarization”, in: Transactions of the Association for Computational Linguistics 12: 39-57.