DH 2026

Daejeon, July 27–31

Fri, July 3109:00–10:30S032103
Long Paper

Authorship Verification of Mihai Eminescu's Feuilletons

Ioana-Roxana Boriceanu
University of Bucharest, Romania · ioana-roxana.boriceanu@s.unibuc.ro
Liviu P. Dinu
University of Bucharest, Romania; Human Language Technologies Research Center · ldinu@fmi.unibuc.ro

Abstract

This study applies an interpretable authorship verification method to Mihai Eminescu's journalistic writings. Using profile-based representations with function words, character n-grams and a rank-based distance, we evaluate disputed texts relative to his verified corpus. Results show that most uncertain articles align stylistically with Eminescu, while a small stable subset falls outside his stylistic range, suggesting different authorship. Qualitative analysis of rejected texts reveals systematic differences in discourse structure and communicative function that distinguish them from the verified corpus.

Introduction and Related Work

Mihai Eminescu is one of the central figures of Romanian cultural history. He is widely regarded as the national poet of Romania and his influence stems from both the originality of his poetry and the range of his intellectual interests. His writings combine philosophical reflection, social critique and political commentary, contributing to his lasting cultural authority (Bolea / Maftei 2025). This variety of interests is especially visible in Eminescu's journalistic activity, produced between 1870 and 1889, which forms a large and influential corpus. Despite its significance, this body of work has received relatively limited sustained attention in scholarship (Mocanu 2020). His articles address politics, education, culture, economics, and social reforms, reflecting the tensions of a society undergoing rapid modernisation (Bitoleanu 2017). Many of these texts were unsigned or published under pseudonyms. Although today we have a fairly reliable picture of the texts attributed to him, a small group of publications still remains of uncertain paternity, making this a particularly relevant setting for authorship verification (AV).

Resolving questions of literary attribution (Juola 2013), authenticating historical manuscripts (Tuccinardi 2017), and analysing disputed texts in forensic settings (Juola 2021) all require methods that can assess whether a document matches an author's known style. Authorship verification (AV) (Stamatatos 2016) addresses this problem directly: rather than selecting among candidate authors as in authorship attribution (Stamatatos 2009), AV asks whether a questioned text is stylistically consistent with a single author's body of work.

Methods for AV span a broad range. Some degrade classification iteratively by removing informative features (Koppel et al. 2007), while others aggregate an author's writings into a single stylometric profile and measure distance to questioned texts (Potha / Stamatatos 2014). Alternatives include compression-based similarity (Halvani et al. 2017) and impostor-based verification, which tests authorship against distractor texts using randomised feature subsets (Koppel / Winter 2014). Recent approaches leverage language models (Kim et al. 2025; Hung et al. 2023), but their opaque representations limit their usefulness in literary and forensic contexts where understanding the basis of a decision matters.

In this work, we apply authorship verification techniques to Eminescu's publicistic corpus to support literary scholarship in resolving texts of uncertain provenance. Our method adapts the profiling approach of Potha / Stamatatos 2014 and integrates the rank distance (Dinu 2003) used with very good results in computational stylistics (Popescu / Dinu 2008; Dinu et al. 2012), producing interpretable stylometric profiles based on function words and character n-grams. We also examine the applicability of such methods to Romanian language data, which remains understudied, and evaluate their suitability for short texts. We have also extended this study in Boriceanu / Dinu 2026, performing a more thorough validation of the proposed method and a deeper philological interpretation of the rejected texts.

Data and Methods

Mihai Eminescu's feuilletons

The corpus is drawn from the Opere edition published by the Romanian Academy (1989), which organises Eminescu's journalism into five chronological volumes. Each volume contains three text categories: published articles (PUB) appearing under his name in contemporary newspapers, manuscripts (MAN) preserved in his handwriting, and texts of uncertain paternity (UPAT) whose attribution remains disputed. Table 1 presents the distribution across volumes.

VolumePUBMANUPAT
Vol. 9 (1870–1877)5105212
Vol. 10 (1877–1878)2641444
Vol. 11 (1880)323117
Vol. 12 (1881)3563114
Vol. 13 (1882–1883, 1888–1889)2429626
Total1695204103

Table 1: Distribution of texts across the five journalistic volumes.

The corpus exhibits substantial variation in document length: texts range from 10 to 12,767 words, with a median of 536 words. Manuscript materials tend to be shorter (median 119 words), while published articles and uncertain texts are considerably longer. For preprocessing, we removed metadata such as headers, editorial notes, and footnotes when present. We applied standard text normalisation, including lowercasing and the removal of artefacts introduced by optical character recognition (OCR). All original punctuation and spacing were preserved, as these contribute to the character n-gram representations.

Methods

Our approach follows a profile-based authorship verification framework inspired by Potha / Stamatatos 2014. All texts with verified authorship (published and manuscript) are treated as positive examples, and the uncertain texts serve as the evaluation set. The author profile is constructed by concatenating all verified texts into a single document, from which feature frequencies are extracted. Each questioned text is represented in the same feature space, and authorship is evaluated by measuring its distance from the profile. We use two feature types: function words (FW) and character n-grams (CNG). For the CNG representation, we retain the 300 most frequent character trigrams in the verified corpus. This fixed-size vocabulary ensures comparability across texts while avoiding sparsity. The function word list (109 items) used in our experiments, presented in Table 2, is based on Dinu et al. 2012 and was slightly adjusted after inspecting the most frequent words in Eminescu's constructed profiles.

și în să se cu o la nu a ce mai din pe un că ca mă fi care era lui fără ne pentru el ar dar îl tot am mi însă într cum când toate al după până decât ei nici numai dacă eu avea fost le sau spre unde unei atunci mea prin ai atât au chiar cine iar noi sunt acum ale are asta cel fie fiind peste această cele face fiecare nimeni încă între aceasta aceea acest aceasta acestei avut ceea cât da făcut noastră poate acestui alte celor cineva către lor unui altă ați dintre doar foarte unor vă aceste astfel avem aveți cei ci deci este suntem va vom vor de cari

Table 2: The function words used in our experiments.

To measure stylistic similarity, we primarily use the rank-based distance presented in Dinu 2003:

Figure 1: The rank-based distance used as the primary stylistic similarity measure.

where fi are the features under comparison and rP(fi) and rQ(fi) denote their ranks in the profile P and the questioned text Q.

For comparison, we also evaluate the asymmetric frequency-based CNG dissimilarity function from Potha / Stamatatos 2014, which serves as the standard baseline for our experiments:

Figure 2: The asymmetric frequency-based CNG dissimilarity function used as a baseline.

where Q is the questioned text, P is the author profile, and fQ(g) and fP(g) are the normalised frequencies of character n-gram g in Q and P, respectively.

For each text of uncertain authorship, we compute its distance to the author profile using both feature sets and both distance formulations. The acceptance threshold is derived from the distribution of distances obtained for the verified Eminescu texts. We use the rule:

Figure 3: Threshold rule used to determine stylistic compatibility.

where μ and σ denote the mean and standard deviation of the distances for genuine texts. We set k = 2, corresponding to approximately 2.3% expected false positives under a normal distribution. This places the threshold well above the typical variation observed around Eminescu's verified writings, accommodating natural variability while still excluding outliers. A questioned text is therefore considered stylistically compatible if its distance satisfies d ≤ T.

Experiments and Results

Experiments

We apply the verification framework to the three corpus subsets: published texts (PUB), manuscripts (MAN), and texts of uncertain paternity (UPAT). We first construct a profile from all published articles and compute CNG and FW distances for all PUB, MAN, and UPAT texts, using thresholds derived from the MAN distribution. To test robustness, we also rebuild the profile from a random subset of the published texts (approximately 80%) together with the manuscripts and estimate thresholds using the remaining subset of published texts.

To assess the contribution of the rank-based distance, we compare it with the asymmetric frequency-based CNG dissimilarity function described in Potha / Stamatatos 2014. We also examined the stability of our method through internal cross-validation, temporal splits, and bootstrap resampling. Across five folds of the MAN set, both feature types achieved near-perfect (95–100%, depending on the split) acceptance of held-out manuscripts, indicating that MAN texts form a compact stylistic cluster. Distances remained stable across volumes and across early and late periods of Eminescu's activity, with no detectable temporal drift. Bootstrap estimates of the MAN-based thresholds showed only limited variation, confirming that the thresholds are robust to sampling differences.

MethodProfileThresholdPUB %MAN %UPAT %
CNG rankPUBMAN97%94%
FW rankPUBMAN100%100%
CNG rank80% PUB + MAN20% PUB98%97%
FW rank80% PUB + MAN20% PUB99%99%
CNG freqPUBMAN100%100%

Table 3: Acceptance rates across feature types and profile configurations.

Results

The acceptance rates obtained for each feature type and profile configuration are presented in Table 3. Overall, the methods behave consistently: both CNG and FW profiles classify the majority of UPAT texts as compatible with the verified corpus, and the outcomes remain stable when the profile is reconstructed from a reduced subset of published texts. This stability indicates that the verification results are not strongly influenced by how the author profile is defined.

Most differences arise from the stricter behaviour of the CNG rank-based distance, which consistently places a small group of UPAT texts just beyond the acceptance threshold. Importantly, the texts rejected in the second profile configuration are a subset of those rejected in the first, indicating stable identification of outliers across settings. A few PUB and MAN texts also fall slightly outside their respective thresholds, which is expected given the statistical derivation of the threshold.

The frequency-based CNG measure, used here as a baseline, yields a very permissive profile, accepting all uncertain texts. Despite this contrast, both distances capture similar stylistic tendencies, with the rank-based version offering sharper discrimination when needed.

TitleVol.WordsSummaryRej.
Fata mamei Ango și Giroflé-Girofla9567theatre commentary with moral reflectionCNG
Nefericitul X9376humorous anecdotal sketchCNG
Ziarele din Viena10475biographical news report on Francisc SchuselkaCNG
Asaltul Angelescu10649political–military critiqueCNG
Din istoria calului10555historical essay on horses in antiquityCNG
O serată literară. Despot Vodă10183cultural news report on a literary salon eventFW
Printr-o indiscrețiune12350diplomatic news brief containing a French telegramCNG

Table 4: UPAT texts rejected by the rank-based classifiers.

Discussion

Table 4 presents the UPAT texts rejected by the stricter rank-based profiles. Examining these cases reveals a shared characteristic: they lack the interpretive or argumentative layer that typifies Eminescu's verified journalism. His authenticated articles consistently move beyond reporting toward causal analysis or institutional critique, whereas the rejected texts remain at the descriptive or anecdotal level.

Nefericitul X exemplifies this pattern: a series of comic vignettes without argumentative progression. Similarly, Din istoria calului compiles historical information without the polemical extension typical of Eminescu's treatment of such material. The case of Printr-o indiscrețiune is distinct: its rejection stems from extended French-language quotations; once removed, the Romanian framing text falls within the acceptance threshold.

We interpret this stable subset of rejections as evidence against Eminescu's authorship. The pattern holds across profile configurations, and the rejected texts share structural properties that distinguish them from his verified work.

Conclusion

This study applied an interpretable authorship verification framework to Mihai Eminescu's journalistic writings, using profile-based representations and rank-based stylistic distances. The results show that most disputed texts fall within Eminescu's stylistic range, while a small stable subset falls outside it, suggesting different authorship. Crucially, the approach provides transparent, reproducible evidence suitable for philological research and close reading. The qualitative analysis of rejected texts reveals systematic differences in discourse structure, demonstrating how verification outcomes can serve as analytic signals that prompt further examination. Future work may incorporate additional stylistic markers, extend the analysis to contemporary authors for contrastive validation, or apply the method to other understudied literary corpora.

Acknowledgements

This research was supported by the Ministry of Education and Research, CNCS-UEFISCDI, project SIROLA, number PN-IV-P1-PCE-2023-1701, within PNCDI IV, and by the project “Romanian Hub for Artificial Intelligence – HRIA”, Smart Growth, Digitization, and Financial Instruments Program, 2021–2027, MySMIS no. 334906.

References
  1. Bitoleanu, Iulian (2017): “The Classicist Vision of the Journalist Mihai Eminescu on Culture”, in: Annals of the University of Craiova for Journalism, Communication and Management 3.1: 95–105.
  2. Bolea, Ștefan / Maftei, Ștefan-Sebastian (2025): “Pessimism, Schopenhauer, and Schopenhauerianism in nineteenth century Romania. The case of the poet Mihai Eminescu”, in: Studies in East European Thought 77.2: 297–311.
  3. Boriceanu, Ioana-Roxana / Dinu, Liviu P. (2026): “Style as Signature: Profile-Based Authorship Verification of Mihai Eminescu’s Journalistic Corpus”, in: Proceedings of the 10th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature 2026: 102–110.
  4. Dinu, Liviu P. (2003): “On the Classification and Aggregation of Hierarchies with Different Constitutive Elements”, in: Fundamenta Informaticae 55.1: 39–50.
  5. Dinu, Liviu P. / Niculae, Vlad / Șulea, Octavia-Maria (2012): “Pastiche detection based on stopword rankings. Exposing impersonators of a Romanian writer”, in: Proceedings of the Workshop on Computational Approaches to Deception Detection: 72–77.
  6. Halvani, Oren / Winter, Christian / Graner, Lukas (2017): “On the usefulness of compression models for authorship verification”, in: Proceedings of the 12th International Conference on Availability, Reliability and Security: 1–10.
  7. Hung, Chia-Yu / Hu, Zhiqiang / Hu, Yujia / Lee, Roy (2023): “Who wrote it and why? Prompting large-language models for authorship verification”, in: Findings of the Association for Computational Linguistics: EMNLP 2023: 14078–14084.
  8. Juola, Patrick (2013): “How a computer program helped show JK Rowling write A Cuckoo’s Calling”, in: Scientific American 20.
  9. Juola, Patrick (2021): “Verifying authorship for forensic purposes: a computational protocol and its validation”, in: Forensic Science International 325.
  10. Kim, Junghwan / Zhang, Haotian / Jurgens, David (2025): “Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains”, in: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: 34855–34880.
  11. Koppel, Moshe / Schler, Jonathan / Bonchek-Dokow, Elisheva (2007): “Measuring differentiability: Unmasking pseudonymous authors”, in: Journal of Machine Learning Research 8.6.
  12. Koppel, Moshe / Winter, Yaron (2014): “Determining if two documents are written by the same author”, in: Journal of the Association for Information Science and Technology 65.1: 178–187.
  13. Mocanu, Mihaela (2020): “Eminescu’s Journalistic Works — Editing Approaches and Reading Patterns”, in: Philobiblon 25.1: 43–61.
  14. Popescu, Marius / Dinu, Liviu P. (2008): “Rank distance as a stylistic similarity”, in: COLING 2008: Companion Volume: Posters: 91–94.
  15. Potha, Nektaria / Stamatatos, Efstathios (2014): “A profile-based method for authorship verification”, in: Hellenic Conference on Artificial Intelligence. Springer: 313–326.
  16. Stamatatos, Efstathios (2009): “A survey of modern authorship attribution methods”, in: Journal of the American Society for Information Science and Technology 60.3: 538–556.
  17. Stamatatos, Efstathios (2016): “Authorship Verification: A Review of Recent Advances”, in: Research in Computing Science 123.1: 9–25.
  18. Tuccinardi, Enrico (2017): “An application of a profile-based method for authorship verification: investigating the authenticity of Pliny the Younger’s letter to Trajan concerning the Christians”, in: Digital Scholarship in the Humanities 32.2: 435–447.