DH 2026

Daejeon, July 27–31

Fri, July 3111:00–12:30S080206-208
Short Paper

The Becoming of an Iconic French Crime Collection: A Computational Analysis of Translated and Original French Le Masque novels 1950-1999

Julia Röttgermann
University of Trier, Germany · roettger@uni-trier.de
Keli Du
University of Trier, Germany · duk@uni-trier.de
Christof Schöch
University of Trier, Germany · schoech@uni-trier.de

Introduction

In 1927, Le Masque, the oldest French crime novel collection, opened with The Murder of Roger Ackroyd by Agatha Christie (Fig 1). In 1930, the collection was institutionally strengthened by the creation of the “Prix du Roman d’Aventures” that aimed to foster the emergence of a French school of detective fiction, despite the influence of British whodunits and American hardboiled crime novels (Martinetti 1997, 19).

Fig. 1 Le Masque book covers 1927, 1929, 1989, 1995.

During the German occupation (1940-1944), books of English or American origin were banned from printing or reprinting in the occupied zone. However, in the aftermath of the liberation and the end of the Second World war, the influence of American literature and culture was strong in France. As French intellectual Jean-Paul Sartre describes it in his famous article “American Novelists in French Eyes” in The Atlantic in 1946: “At once, for thousands of young intellectuals, the American novel took its place, together with jazz and the movies, among the best of the importations from the United States“ (Sartre 1946). At the same time, alongside the translated novels from the United States and the United Kingdom, the proportion of original French crime novels within the Le Masque crime series grew steadily over the decades (Fig. 2).

Fig. 2 Number of published Le Masque novels per decade.

Today, Le Masque has come to epitomize the development of the genre in France and is still in continuous publication (Reuter 2001). The present study traces genre formation in Le Masque through contrastive analysis of translations versus French originals, examining the unique characteristics of each text group.

Previous Approaches to distinctiveness in crime fiction

From a literary studies point of view, a multitude of works deal with recurring narrative patterns within crime novels (e.g. Todorov 1971, Boileau-Narcejac 1982, Lits 1999). Concerning distinctive words in French and English crime novels, Novakova and Gymnich (2021) have demonstrated a statistically relevant overrepresentation of certain linguistic phenomena in crime fiction, such as the motive découvrir le corps / find the body, as have Gonon et al. (2018) for the motive la porte / the door. In the Digital Humanities, Barré et al. (2025) investigate the vocabulary that is the most typical of certain detective archetypes over time. Kraif (2017) compares textual statistics between original and translated-to-French novels. To the best of our knowledge, the crime series Le Masque has not yet been the object of a computational distinctiveness analysis.

Data

Our main corpus consists of 265 novels published between 1950 and 1999 in Le Masque, including 134 translated and 131 original French novels (Fig 3). The translations originate mainly from the UK and USA.

Fig. 3 Number of Le Masque novels per decade in our corpus.

We also use an additional corpus of 392 French crime novels (published between 1950 and 2009) to ensure sufficient data for training topic models and to explore the distinctive characteristics of Le Masque novels.

Methods

Stylometric analysis and distinctive words: First, we examined our premise of a difference between the original French Le Masque novels and the translated ones in two preliminary tests: (1) stylometric analysis with stylo (Eder et al. 2016) using 1000 Most-Frequent-Words and Cosine Delta (Evert et al. 2017), and (2) identifying top distinctive words of each part using logarithmic Zeta, which has been shown to extract meaningful content words (Röttgermann et al. 2025).

Distinctive topics: Secondly, we analyzed distinct topics within each text group using topic modeling. For this, all novels were segmented in 1000-word-chunks and a list of stop words (including function words, common French and English person names) was removed from the corpus, and lemmatization was performed. We trained three topic models using the same parameter settings but different random seeds.

We used MALLET (McCallum 2002) with alpha=1, beta=0.005, num-iterations=1500, optimize-burn-in=200, optimize-interval=200. We trained models with 20, 50, 100 topics and decided for 100 topics, which best captures the granularity of topic distinctiveness in our corpus.

To identify distinctive topics of each text group, we used Welch’s T-Test (cf. Welch 1947, Du et al. 2026). Across the three models, we observed topics that refer to the same themes and consistently show significant differences between the groups. These recurring topics best illustrate certain differences between the text groups.

Results

Stylometric differences between translated and non-translated Le Masque novels

Fig. 4 Stylometric Study, detailed view on a cluster with novels by different authors all translated by Marie-Louise Navarro: Le Petit Été de la Saint Luke (1980), Le Trèfle à cinq feuilles (1975), Les Arbres de Josué (1973), Tornade sur la ville (1981), Une femme très douce (1969).

We used stylo in R for a cluster analysis and observed two stylistic groups (translated, non-translated; Fig. 4). Furthermore, there are groupings visible related to stylistic traces of specific translators, such as Marie-Louise Navarro, with all novels translated by her clustering on the same branch, showing a strong ‘translator’s signal’.

On the stylometric (in)visibility of translation, see Rybicki 2012.

The opposite, a branch of novels by John Dickson Carr, translated by three different translators nevertheless clustering together, can also be observed, indicating that the authorial style is more dominant than the translational style, in this particular case.

Lexical distinctiveness with Zeta

For the translated novels, the top distinctive words are mostly English titles (‘Mr’, ‘Mrs’,‘Miss’) and words indicating English settings (‘London’) or architecture (‘cottage’), as well as typical investigators (‘Scotland Yard’, ‘detective’). By contrast, original French novels exhibit distinctive vocabulary tied to French institutional (‘Franc’), investigatory (‘commissioners’), and urban contexts (‘platform’). Besides the frequency of English place names or the mention of European currency, there are more subtle differences, such as an emphasis on fear (‘anxiety’) and the victim (‘victim’) in French crime novels.

Lexical distinctiveness with Zeta results: https://github.com/Zeta-and-Company/LeMasque/blob/main/DH2026/265_novels/zeta_results.csv.

This encouraged us to further explore the differences not only at the word level, but also using more complex and semantically-richer linguistic features, namely topics.

Distinctive topics for Le Masque

We first contrasted all Le Masque novels with other French crime novels in our corpus to identify the thematic specificities of Le Masque. A striking result was that other crime collections show topics containing sociolect and oral language that appear to be present to a lesser degree in Le Masque. These distinctive topics are dominated by colloquial markers such as type, flic, gars, quoi, hinting at the depiction of a lower-class milieu and spoken immediacy.

See topic_81, topic_65, topic_79: https://github.com/Zeta-and-Company/LeMasque/blob/main/DH2026/265_novels/distinctive_topics_LeMasque-NonLeMasque.csv.

Furthermore, contrastive topics emerge for the character of Maigret, a famous detective character. Even though Le Masque novels also contain stylistic traces of orality and slang, other French crime series like Série noire contain much more such colloquial markers, when considered in a contrastive perspective.

This quote by author Thomas Narcejac 1954 on the specificity of Le Masque in contrast to other French crime series underscores the gentleman-like tonality: “Le Masque ressemble à Agatha Christie. Le grand mérite de cette collection, c’est d’avoir toujours donné à l’énigme sa première place.” (Martinetti 1997, pp. 20)

Distinctive topics for translated Le Masque novels

The topics that are distinctive for the translated Le Masque novels (that is, over-represented in the translations when compared to the French originals) reveal deeper insights into the specificities of this group of novels: topic 49 revolves around family lines (‘aunt’, ‘uncle’, ‘father’, ‘cousin’) and topic 7 around inheritance (‘will’, ‘notary’, ‘inherit’). Both topics hint at narrative conventions of British whodunits, where legal-familial tensions often play a role in inheritance mysteries. These topics suggest that translated novels import certain modes of crime and investigation into the French crime novel series.

Distinctive topics for original French Le Masque novels

Concerning the original French novels, a distinctive topic points to travel narratives (topic 24: ‘plane’, ‘french’, ‘hour’, ‘hotel’, ‘girl’, ‘american’), containing the top topic word ‘american’ which hints at the preoccupation of French authors with American and English settings and figures, even in novels written in French. French authors of the early decades of Le Masque series often use pseudonyms and adopt an American persona, also imitating the style of the American hardboiled novel in using slang and oral language.

An example is Marc Demeulenaere, who used the English pseudonym “Mark Demwell” to publish Le Crime des hauts-toupets (1979), Le Trésor d'Irbla (1980) and Meurtre au collier (1980).

Topic 55 is proof of this slang-laden or oral style in original French Le Masque novels (‘guy’, ‘yes’, ‘dude’, ‘cop’, ‘what’).

Natacha Levet (2024, 2) describes this “American moment” in French crime fiction: “le roman noir à l’américaine va devenir un polar à la française, qui hérite aussi bien de la langue verte des romans de truands de l’entre-deux guerres que de la stylistique du roman prolétarien.”

The topic is also interwoven with top topic words associated with the character of the inspector, typical for French crime fiction (topic 55: ‘inspector’, ‘cop’). The inspector or commissioner (topic topic 31: ‘commissioner’, ‘Monsieur’, ‘inspector’) represents a distinctive French type of investigator, being represented for example by ‘Danglard’ and ‘Adamsberg’, recurring characters in Fred Vargas’ crime novel series.

On Adamsberg as an antithesis of the “supermen detectives of the past” cf. Platten 2011, 221-250. On characteristic detective types per time cluster in crime fiction cf. Barré et al. 2025.

In addition to the topics that are distinctive for each of the two groups, there are also some topics present that are prominent within the Le Masque group as a whole, without being distinctive by comparison with the other crime fiction novels, such as topics on crime scenes (topic 12: ‘body’, ‘corpse’, ‘doctor’, ‘trace’, ‘find’), alcohol consumption (topic 79: ‘glass’, ‘drink’, ‘bottle’, ‘whiskey’), court hearings (topic 65: ‘lawyer’, ‘witness’, ‘judge’, ‘jury’) or funerals (topic 64: ‘church’, ‘cemetery’, ‘coffin’).

Conclusion

Overall, we have demonstrated that combining topic modeling, stylometry and statistical measures of distinctiveness is useful and can effectively capture meaningful genre-specific features. Our analysis shows that Le Masque became an iconic crime collection not simply by importing Anglo-American detective fiction, but by transforming it into a distinct French form. Translated novels retain features of British and American crime fiction: inheritance plots, family constellations, and English settings, whereas original French novels show local markers such as French institutions, the figure of the inspector or commissioner, as well as an emphasis on the victim and anxiety as a contrastive element. At the same time, French authors appropriated foreign models through American settings or markers of orality and hardboiled slang. Le Masque thus became canonical through this negotiation between translation and invention: imported genre conventions were reshaped into a recognizable French crime-fiction identity.

Data Availability

Data can be found here: https://doi.org/10.5281/zenodo.20054189.

References
  1. Barré, Jean, Olga Seminck, Antoine Bourgois, and Thierry Poibeau. 2025. “Modeling the Construction of a Literary Archetype: The Case of the Detective Figure in French Literature”. Anthology of Computers and the Humanities 3: Computational Humanities Research 2025: 983–99. 10.63744/SMbYIWcHZj87.
  2. Boileau-Narcejac. 1982. Le roman policier. Que sais-je ? Presses Universitaires de France.
  3. Du, K., Röttgermann, J., & Schöch, C. 2026. “Keyness Measures und BERTopic kombiniert: Eine Distinktivitätsanalyse von Subgenres des französischen Romans”. DHd 2026 Nicht nur Text, nicht nur Daten (DHd2026), Wien, Österreich. https://doi.org/10.5281/zenodo.18696363.
  4. Eder, Maciej, Jan Rybicki, and Mike Kestemont. 2016. “Stylometry with R: A Package for Computational Text Analysis”. The R Journal 8 (1): 107–21. 10.32614/RJ-2016-007.
  5. Evert, Stefan, Fotis Jannidis, Thomas Proisl, et al. 2017. “Understanding and Explaining Distance Measures for Authorship Attribution.” Digital Scholarship in the Humanities. 10.1093/llc/fqx023.
  6. Gonon, Laetitia, Vannina Goossens, Olivier Kraif, Iva Novakova, und Julie Sorba. 2018. “Motifs textuels spécifiques au genre policier et à la littérature blanche”. SHS Web of Conferences 46: 06007. https://doi.org/10.1051/shsconf/20184606007.
  7. Kraif, Olivier. 2017. “Traduire le polar : une étude textométrique comparée de la phraséologie du roman policier en français source et cible”. Synergie Pologne, De la phraséologie aux genres textuels : état des recherches et perspectives méthodologiques, Numéro 14. https://gerflint.fr/Base/Pologne14/pologne14.html.
  8. Levet, Natacha. 2024. “Du polar « amerloque » au roman de langue verte : enjeux génériques du style « roman noir »”. Acta fabula. https://www.fabula.org:443/colloques/document12580.php.
  9. Lits, Marc.1999. Le roman policier: introduction à la théorie et à l’histoire d’un genre littéraire. Editions du CEFAL.
  10. Martinetti, Anne. 1997. Le Masque: histoire d’une collection. Encrage.
  11. McCallum, Andrew Kachites. 2002. “MALLET: A Machine Learning for Language Toolkit.” http://mallet.cs.umass.edu.
  12. Novakova, Iva, und Marion Gymnich. 2021. “Extended Phraseological Units and Literary Genres: A Contrastive Analysis of French and English Lexico-Syntactic Constructions”. Lexique, Nr. 28: 87–112.
  13. Platten, David. 2011.“Mapping Minds and Figuring Plots: The Novels of Fred Vargas”. In The Pleasures of Crime: Reading Modern French Crime Fiction. Rodopi.
  14. Reuter, Yves. 2001. Le roman policier. Nathan.
  15. Röttgermann, Julia, Keli Du, Julia Havrylash, and Christof Schöch. 2025. “Expertise vs. Statistics. A Qualitative Evaluation of Three Keyness Measures (Logarithmic Zeta, Welch’s t-Test, and Log-Likelihood Ratio Test) Applied to Subgenres of the French Novel”. Digital Humanities Quarterly 19 (3). https://dhq.digitalhumanities.org/vol/19/3/000816/000816.html.
  16. Rybicki, Jan. 2012. “The Great Mystery of the (Almost) Invisible Translator”. In Quantitative Methods in Corpus-Based Translation Studies: A Practical Guide to Descriptive Translation Research. John Benjamins. 10.1075/scl.51.09ryb.
  17. Sartre, Jean-Paul. 1946. “American Novelists in French Eyes”. The Atlantic.
  18. Tzvetan Todorov. 1998. “Typologie des Kriminalromans”, in: Jochen Vogt (Hrsg.), Der Kriminalroman. Poetik. Theorie. Geschichte, München: Wilhelm Fink, S. 208-215.
  19. Welch, Bernard Lewis. 1947. “The Generalization of Student’s Problem When Several Different Population Variances Are Involved”. Biometrika 34 (1–2): 28–35. 10.1093/biomet/34.1-2.28.