DH 2026

Daejeon, July 27–31

Thu, July 3016:30–18:00S067106
Short Paper

The People Without a Word: Tracing the Fragmented History of Gender Nonconformity in British Newspapers

Jiwoo Choi
KAIST, Korea, Republic of (South Korea) · jwchoi0515@kaist.ac.kr
Seohyon Jung
KAIST, Korea, Republic of (South Korea) · seohyon.jung@kaist.ac.kr

Abstract

This study examines how the amorphous and collective concept of "gender-nonconforming" has been socially constructed and differentiated over time, using a large-scale text analysis of a historical British newspaper dataset. We argue that contemporary identity terms such as "transgender," "nonbinary," and "butch" are not intrinsic entities but rather outcomes produced and shaped within historical discourses and institutional classification systems. Thus, the research goal is not to "discover modern identities" within past newspaper records, but to trace how the concept of gender-nonconforming as a gelatinous mass resisting definition has been cut, named, and reassembled by linguistic and social forces. To achieve this, we apply NLP-based analysis techniques to detect discourse shifts by period, region, and newspaper type, and read closely the articles in which such shifts come into view. Furthermore, by examining how these structures represent or challenge contemporary AI-based text classification systems, it offers new insights into digital archives and infrastructures of cultural memory.

Introduction & Background

Gender-nonconforming bodies have often been positioned on the periphery of institutional records and cultural memory. In official materials such as newspapers, legal documents, medical reports, and police records, they have been represented in various forms, and social knowledge about them, such as their gender or sexuality, has been produced and circulated in specific ways. Scandal, crime, pathology, ridicule. They have existed as a loose, fluid, spectral spectacle, constantly being joined, divided, and rearranged within different meanings, emotions, and power relations. However, existing research has tended either to project modern identity categories onto the past (Halberstam 1998; Stryker 2008) or to remain confined to discovering queer history in newspaper articles or archives (Cocks 2003; Houlbrook 2005). But how has society handled the mass of gender nonconformity? How have these loose, fluid, spectral effects that have been joined, divided, and rearranged at each specific discursive moment existed within society? What this research intends to do is not to find modern concepts within past newspapers, but rather to trace the process by which beings within the larger mass were dissected, sutured, and then erased again under the scalpel of language.

Particularly, as the newspaper medium both shaped public discourse and was deeply intertwined with state, medical, legal, and religious power, the ways it represented gender nonconformity should be read as a significant case revealing the politics of classification, naming, and framing, going beyond simple article production. Today, large-scale newspaper archives combined with OCR and NLP technologies possess the capacity to reopen this inquiry (Beelen et al. 2023; Hamilton et al. 2016). This study traces how gender-nonconforming bodies were classified, named, and framed, and how the landscape of this discourse diverged across time, region, and newspaper type. This is not a restorationist study seeking to discover "trans" or "queer" within newspaper archives. Instead, it analyzes how gender-nonconforming bodies were constructed, consumed, and erased through linguistic devices and classification systems within the discursive machinery of newspapers.

Using the NewsWords dataset (Beelen 2025), which includes a wide range of titles from British newspapers across the long nineteenth century, we analyze how the amorphous, collective concept of "gender-nonconforming" was socially constructed and differentiated. This analysis aims to trace how gender-nonconforming bodies were named, categorized, and framed within diverse genre-based and affective frames, including criminalization, pathologization, ridicule, scandal, and human rights discourse. Ultimately, it seeks to reveal how discursive devices have constructed and erased gender nonconformity, going beyond a restorationist approach that projects contemporary identity categories into the past.

Research Questions

RQ1. How were they talked about? What linguistic and emotional frames constituted the collective mass of "gender-nonconforming" within nineteenth-century British newspaper discourse?

RQ2. How did they change over time? How did naming conventions and classification systems indicating gender nonconformity diverge and reorganize according to period, region, and newspaper type?

RQ3. What did those discourses do? How was the spectrum of gender nonconformity segmented, rearranged, and erased within newspaper discourse?

This study begins by understanding gender nonconformity not as a pre-given identity or pre-existing classification, but as an effect constituted within historical discourses and institutions. As Butler (2002) states, "There is no gender identity behind the expressions of gender." This study analyzes "gender-nonconforming" not as a fixed category but as a collective effect repeatedly reconfigured within diverse frames such as criminalization, pathologization, caricature, scandal, and rights discourse. While this study does not view newspapers as singular meaning-producing entities, it posits that newspapers function as crucial mediating devices where discourses on crime, medicine, religion, and law intersect, thus playing a central role in the construction and circulation of social meanings surrounding sexuality and gender (Foucault 1990; Bowker / Star 2000; Sedgwick 1990).

Methodology & Data

By treating “gender-nonconforming” as a gelatinous mass, we mean a semantic field whose neighbourhood structure in diachronic embeddings does not settle into discrete clusters; it is this instability, rather than any underlying referent it might be made to disclose, that our method takes as its object. This study draws on the NewsWords dataset (Beelen 2025), which provides monthly word counts from British newspapers digitized by the British Library across the long nineteenth century (1780–1920). The contextualized release links these counts to publisher metadata derived from Mitchell's Newspaper Press Directories (1846–1920), recording political leaning, price, and place of publication alongside roughly 117 billion tokens drawn from a vocabulary of 196,719 unique terms. That the corpus overrepresents mainstream conservative and liberal titles (Beelen et al. 2023) is itself a condition of the inquiry rather than a flaw to be corrected; the dominant press is precisely the apparatus through which gender-nonconforming bodies were rendered legible to the reading public, and it is that machinery of public naming, rather than any prior interiority beneath it, that this study takes as its object.

Using this corpus, we perform NLP-based computational analysis. This study does not presuppose a fixed list of gender-nonconforming vocabulary. Instead, it begins from a multi-layered seed lexicon distributed across three discursive registers, medical-sexological, legal-criminal, and vernacular-derogatory, drawn from period dictionaries, the OED Historical Thesaurus, and prior historical scholarship (Cocks 2003; Houlbrook 2005; Crozier 2008). These seeds are entry points into distinct discursive registers, not referents to a stable identity behind them. From these seeds, we extract neighbouring words from diachronic word embeddings and broadly collect tentative candidate expressions through co-occurrence analysis at the newspaper-month level, drawing on methods developed for historical text (Hamilton et al. 2016; Garg et al. 2018; Kozlowski et al. 2019). These expressions are clustered by their usage contexts to construct a discursive repertoire. This enables a macro-level understanding of how the semantic field and framing modes of gender nonconformity were constituted and transformed.

The close reading takes as its subject the representative articles and semantic networks derived from the computational analysis results. Articles are sampled at points where the macro analysis registers reorganisation, shifts in vocabulary, in register, in metadata distribution, and read against the patterns from which they were drawn. In this process, we analyze linguistic framing devices, affective expressions, and power relations, interpreting how the spectrum of gender nonconformity was segmented, hierarchised, and erased within discourse. This allows us to analyze the social and historical context in which the discourse operated, rather than merely identifying statistical patterns.

In revealing how gender-nonconforming subjects were treated differently within diverse frames of criminalization, pathologization, ridicule, scandal, and rights discourse, we trace how this undefined mass of gender nonconformity was segmented and formed under different discourses and social pressures throughout the nineteenth century.

References
  1. Beelen, Kaspar (2025): NewsWords Data (Contextualized Word Counts). Zenodo. DOI: 10.5281/zenodo.14961220 [08.05.2026].
  2. Beelen, Kaspar / Lawrence, Jon / Wilson, Daniel C. / Beavan, David (2023): "Bias and representativeness in digitized newspaper collections: Introducing the environmental scan", in: Digital Scholarship in the Humanities 38, 1: 1–22.
  3. Bowker, Geoffrey C. / Star, Susan Leigh (2000): Sorting Things Out: Classification and Its Consequences. Cambridge, MA: MIT Press.
  4. Butler, Judith (2002): Gender Trouble. New York: Routledge.
  5. Cocks, Harry G. (2003): Nameless Offences: Homosexual Desire in the Nineteenth Century. London: I. B. Tauris.
  6. Crozier, Ivan (2008): Sexual Inversion: A Critical Edition. Basingstoke: Palgrave Macmillan.
  7. Foucault, Michel (1990): The History of Sexuality: An Introduction. New York: Vintage.
  8. Garg, Nikhil / Schiebinger, Londa / Jurafsky, Dan / Zou, James (2018): "Word embeddings quantify 100 years of gender and ethnic stereotypes", in: Proceedings of the National Academy of Sciences 115, 16: E3635–E3644.
  9. Halberstam, Judith (1998): Female Masculinity. Durham, NC: Duke University Press.
  10. Hamilton, William L. / Leskovec, Jure / Jurafsky, Dan (2016): "Diachronic word embeddings reveal statistical laws of semantic change", in: Association for Computational Linguistics (ed.): Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Berlin, August 2016: 1489–1501.
  11. Houlbrook, Matt (2005): Queer London: Perils and Pleasures in the Sexual Metropolis, 1918–1957. Chicago: University of Chicago Press.
  12. Kozlowski, Austin C. / Taddy, Matt / Evans, James A. (2019): "The geometry of culture: Analyzing the meanings of class through word embeddings", in: American Sociological Review 84, 5: 905–949.
  13. Ryan, Yann / McKernan, Luke (2021): "Converting the British Library's Catalogue of British and Irish Newspapers into a public domain dataset: Processes and applications", in: Journal of Open Humanities Data 7, 0: 1. DOI: 10.5334/johd.23.
  14. Sedgwick, Eve Kosofsky (1990): Epistemology of the Closet. Berkeley: University of California Press.
  15. Stryker, Susan (2008): Transgender History. Berkeley: Seal Press.