Daejeon, July 27–31
This article focuses on both human and AI abilities to pastiche literary works written in late 19th century Romanian. This endeavor is typical for DH, being a challenging task since Romanian is a low resourced language and its archaic version is not standardized. Pastiche is a stylistic imitation, acknowledging the source and often paying it a tribute, without deceptive intent. There are several types of literary pastiche, each focusing on different aspects of the original work, like character pastiche, style pastiche, idea pastiche, genre pastiche, or combined pastiche which incorporates elements from multiple sources or ideas from different works.
In this work we will explore the style pastiche, since we are interested in imitating human stylome, in the sense of the author’s stylistic signature.
Our research questions are:
RQ1. Can we quantify the quality of a literary pastiche? and
RQ2. Can AI pastiche literary works comparable to professional writers?
To answer these fundamental questions, we present a case study of a Romanian 19th century author, Mateiu Caragiale (1885–1936), that has been pastiched by another professional writer, Radu Albala (1924-1994) at the beginning of the 20th century. Mateiu Caragiale was one of most important Romanian authors who contributed to the modernization of the Romanian literary language, through atmospheric novels written with a rich, poetic vocabulary. He left unfinished his last novel Sub pecetea tainei [Under the seal of secrecy] and Radu Albala continued it with the novel În deal pe Militari [On the Militari hill] with the declared purpose of imitating the original author’s style. Some other Romanian writers were inspired by him, like Ion Iovan, who authored a Mateiu Caragiale diary called Ultimele însemnări ale lui Mateiu Caragilale [Last Notes of Mateiu Caragiale], or the contemporary writer Ștefan Agopian, who did not attempt to pastiche Mateiu’s work, but is considered one of his closest stylistic successors.
Previous studies investigated the possibility of clustering and automatically classifying Mateiu and Albala’s novels (Dinu et al. 2008), showing that the classical machine learning algorithms like SVMs correctly clustered Mateiu’s and Albala’s novels, but only by a narrow margin. The study showcased the importance of punctuation and stop words in authorship attribution. A subsequent study (Dinu et al. 2012) included in the dataset Ion Iovan and Ștefan Agopian, showing that the clustering dendograms of the novels written by all four authors were closer to the reality if Rank Distance (Dinu 2005) rather than Euclidian distance was used.
Two recent articles introduced a dataset with LLM-generated pastiches (Dinu et al. 2025a; Dinu et al. 2025b) after Mateiu’s unfinished novel. They showed that LLMs can write fairly good quality pastiches, measured by standard quality metrics, but still not as good as the professional writers.
In this study, we combine all previous datasets and add pastiches generated by the latest LLMs. We perform an in-depth analysis on this dataset by extracting psycho-linguistic features with LIWC tool (Boyd et al. 2022) and Romanian dictionary (Dudău and Sava 2020) and computing Language Style Matching scores between all authors, by performing sentiment analysis and Principal Component Analysis (PCA) visualization on 23 stylistic features.
The corpus contains 6 chapters by Mateiu Caragiale, plus the unfinished novel, 6 novels by Radu Albala, plus his continuation of Mateiu’s unfinished novel, 6 texts by Ion Iovan, 6 works by Ștefan Agopian, 6 LLM-generated pastiches in December 2024, and 6 LLM-generated pastiches generated in January 2026.
We extracted the psycho-linguistic features on the basis on which LIWC tool computed the Language Style Matching (LSM) scores between any of the 6 authors (4 human writers plus old and new LLMs).
Figure 1. Heatmap of LSM
The LSM scores listed in Figure1 show that the human writers resemble each other more than they resemble the original work by Mateiu. Agopian is the closest to Mateiu’s style according to the LSM score. The 2024 LLMs have the lowest scores with every other category, with only 0.32 LSM score with Mateiu, while the 2026 LLMs score higher than Albala and Iovan (0.69 vs. 0.67) and are remarkably closer to Mateiu than the 2024 versions (0.69), suggesting the newer models produce “better” pastiches. The closest LSM score to Mateiu was obtained by Agopian, 0.79, much higher than all the others, confirming he is one of the writers closest to Mateiu Caragiale’s style.
The sentiment analysis performed with Romanian BERT
https://huggingface.co/readerbench/ro-sentiment
Figure 2. Sentiment scores for the six authors
Finally, we extracted 23 stylometric features: words per sentence, type-token ratio, nominal/ adjective/adverb/pronoun ratio, sentence length, hapax ratio (words appearing only once), sentiment score (Romanian BERT), pacing (comma/dot ratio), 13 punctuation and function word features (comma, dot, exclamation, question mark, semicolon, colon, dash, parenthesis, quote, conjunction, preposition, article, pronoun). We applied the Principal Component Analysis (PCA) visualization technique to reduce these distinctive stylistic features to only two printable dimensions, obtaining the visual representation in Figure 3. In this representation, Albala is again the closest to Mateiu, shortly followed by 2026 LLMs, which seem to have improved again considerably from 2024. Based on this representation, we made the correct prediction that the novel Sub pecetea tainei by Mateiu Caragiale is attributed to Mateiu, and its continuation În deal pe Militari is attributed to Albala, based on the shortest distance to a centroid. This suggests that there are powerful stylistic features that distinguish between the authors.
Figure 3. PCA visualization of a full set of stylistic features for all six authors
Overall, the Romanian professional writers from the 20th century are still closer in style to Mateiu’s work than the LLMs, but the gap shrinks fast.
This work was supported by a grant of the Ministry of Education and Research, CCCDI - UEFISCDI, project number PN-IV-P6-6.1-CoEx-2024-0042, within PNCDI IV, and by a grant of the Ministry of Research, Innovation and Digitization, CNCS - UEFISCDI, project SIROLA, number PN-IV-P1-PCE-2023-1701, within PNCDI IV.