DH 2026

Daejeon, July 27–31

Short Paper

FROM GENES TO TOKENS: A GWAS INSPIRED APPROACH FOR INTERPRETABLE STYLOMETRIC ANALYSIS The study was implemented in the framework of the Basic Research Program at HSE University HSE-BR-2025-031

Dmitry Pronin
HSE, Russian Federation · ddpronin003@yandex.ru
Evgeny Kazartsev
HSE, Russian Federation · kazar@list.ru

Introduction

Computational stylometry is a field of Digital Humanities concerned with authorship attribution through the quantitative analysis of linguistic features. Its central assumption is that authors exhibit relatively stable and partly unconscious stylistic habits that can be detected statistically across texts. Although recent transformer-based methods often achieve high predictive accuracy, they usually do so at the cost of interpretability (Huang et al. 2025). Yet interpretability remains a central concern in stylometric research, from classical authorship-attribution studies to more recent work on characteristic features and explainable Delta-like measures (Stamatatos 2009; Burrows 2002; Evert et al. 2017; Klaussner et al. 2015; Savoy 2012).

This paper proposes a complementary perspective inspired by genome-wide association studies (GWAS). In classical GWAS, each single nucleotide polymorphism is tested for association with a phenotype such as disease status or a quantitative trait. A separate regression model is fitted for each variant, its effect size and p-value are estimated, and significance thresholds are corrected to control for multiple testing (Uffelmann et al. 2021; Visscher et al. 2017). Transferred to stylometry, this logic makes it possible to move from a single aggregate distance between authors toward a feature-wise map of statistically interpretable lexical contrasts.

The proposed approach therefore does not replace existing stylometric distance measures. Instead, it adds an inferential layer that allows individual lexical features to be tested directly for their association with authorship. This makes it possible to identify statistically significant markers and to interpret stylistic contrasts at the token level.

Stylometric «GWAS»

A stylometric analogue of GWAS can be formulated by treating standardized token frequencies as “genes” and authorship as the “phenotype”. Whereas classical GWAS usually deals with binary or quantitative traits, the same logic can be extended to multiclass attribution through one-vs-rest decomposition. In contrast to Burrows’s Delta, which compresses deviations across many frequent words into a single distance (Burrows 2002; Evert et al. 2017), the proposed approach keeps inference at the level of individual tokens. As a result, each word or n-gram receives its own statistically interpretable contribution to the distinction between authors.

For each token, a separate univariate regression is fitted in which the binary authorship label is predicted from the standardized frequency of that token in a chunk. The regression coefficient estimates the direction and strength of association with the target author, while the corresponding p-value tests whether this association differs significantly from zero.

Because the number of tested features may reach several thousand, significance thresholds must be corrected in order to limit false positives. In GWAS this is commonly addressed through family-wise error or false-discovery-rate control; in the present study Bonferroni correction is used as the most transparent baseline (Uffelmann et al. 2021).

Figure 1. Stylometric “GWAS” pipeline.

Experiments and results

The method was tested on English, German, and Russian corpora containing texts by three authors. Each corpus comprised 75 documents, with three texts per author. After lemmatization, the texts were segmented into non-overlapping chunks of 10,000 lemmas, and the 5,000 most frequent lemmas in each corpus were retained for analysis. A logistic-regression model was fitted for every unigram, after which Bonferroni correction was applied to the resulting p-values.

H. G. Wells.

For H. G. Wells, the procedure identified 221 significant unigrams associated with his distinctive style. The strongest markers combine high-frequency function words with semantically meaningful items such as empire, tremendous, and anyhow, suggesting a profile that blends narrative scale, evaluative intensity, and conversational modulation. On the Figure 2 green points represent unigrams with positive regression coefficients while red points indicate unigrams with negative coefficients. Green unigrams are common for the author whereas red ones are notably avoided in his works.

Figure 2. Manhattan plot for H. G. Wells.

H. Hesse.

In the German corpus, statistically significant lemmas specific to Hermann Hesse emphasize an introspective and philosophical register (Figure 3). Words such as angst (“fear”), erinnerung (“memory”), wahnsinnig (“mad”), and hinweg (“away, beyond”) point to inner experience, self-reflection, and emotional intensity. Other markers, including spiegeln (“to reflect”), still (“quiet”), and roch (“smelled”), indicate a fine-grained attention to perception and subtle existential states.

Figure 3. Manhattan plot for H. Hesse.

L. N. Tolstoy.

In the Russian corpus, the analysis revealed a broad set of significant lemmas associated with Leo Tolstoy’s style (Figure 4). Markers such as сказать (“to say”), чувствовать (“to feel”), увидать (“to see”), and услышать (“to hear”) indicate a narrative mode grounded in perception and reporting. Other recurrent features, including самый (“the most”), очевидный (“obvious”), and несмотря (“despite”), reflect emphasis, evaluation, and syntactic structuring. Taken together, these signals are consistent with Tolstoy’s realist style, marked by detailed observation and discursive narration.

Figure 4. Manhattan plot for Leo Tolstoy.

For all authors, permutation testing showed that the observed number of significant features exceeded the number expected under random labeling. A public code repository containing preprocessing scripts, regression procedures, multiple-testing correction, and Manhattan-plot visualization is being prepared to support reproducibility.

GitHub repository: https://github.com/DDPronin/GWAS-stylometry

Future Development

Further development naturally concerns the introduction of stylometric covariates. In GWAS, variables such as age, sex, or principal components of population structure are routinely added to the regression model in order to reduce confounding and prevent spurious associations. In stylometry, an analogous role may be played by genre, time period, topic, and other linguistic or cultural parameters. Incorporating such covariates would help distinguish genuine authorial markers from features driven by external context. At the same time, Manhattan plots make it possible to organize detected signals not only by frequency, as in the present study, but also by predefined linguistic groupings, for example parts of speech or semantic classes.

Conclusion

This paper presents an interpretable approach to authorial-style analysis based on the logic of GWAS. Each token is evaluated independently for its association with authorship, which makes it possible to identify statistically significant lexical markers rather than relying only on aggregate distances. Across English, German, and Russian corpora, the method recovers plausible and linguistically meaningful stylistic signals.

References
  1. Burrows, John F. (2002): “‘Delta’: A measure of stylistic difference and a guide to likely authorship”, in: Literary and Linguistic Computing 17, 3: 267–287.
  2. Evert, Stefan / Proisl, Thomas / Jannidis, Fotis / Reger, Isabella / Pielström, Steffen / Schöch, Christof / Vitt, Thorsten (2017): “Understanding and explaining Delta measures for authorship attribution”, in: Digital Scholarship in the Humanities 32, suppl. 2: ii4–ii16.
  3. Baixiang Huang / Canyu Chen / Kai Shu, Kai (2025): “Authorship Attribution in the Era of LLMs: Problems, Methodologies, and Challenges”, in: ACM SIGKDD Explorations Newsletter 26, 2: 21–43.
  4. Klaussner, Christoph / Nerbonne, John / Çöltekin, Çağrı (2015): “Finding Characteristic Features in Stylometric Analysis”, in: Digital Scholarship in the Humanities 30, suppl. 1: i114–i129.
  5. Savoy, Jacques (2012): “Authorship Attribution Based on Specific Vocabulary”, in: ACM Transactions on Information Systems 30, 2: Article 12, 1–30. DOI: 10.1145/2180868.2180874
  6. Stamatatos, Efstathios (2009): “A survey of modern authorship attribution methods”, in: Journal of the American Society for Information Science and Technology 60, 3: 538–556. DOI: 10.1002/asi.21001.
  7. Uffelmann, Emil / Huang, Qin Qin / Munung, Nchangwi Syntia / de Vries, Jantina / Okada, Yukinori / Martin, Alicia R. / Martin, Hilary C. / Lappalainen, Tuuli / Posthuma, Danielle (2021): “Genome-wide association studies”, in: Nature Reviews Methods Primers 1: 59.
  8. Visscher, Peter M. / Wray, Naomi R. / Zhang, Qi / Sklar, Pamela / McCarthy, Mark I. / Brown, Matthew A. / Yang, Jian (2017): “10 years of GWAS discovery: biology, function, and translation”, in: The American Journal of Human Genetics 101, 1: 5–22.