DH 2026

Daejeon, July 27–31

Thu, July 3011:00–12:30S093204-205
Long Paper

Collapsing Probability Space: Rethinking Representationalism with BERT

Katherine Bode
Australian National University, Australia · katherine.bode@anu.edu.au
Galen Cuthbertson
Australian National University, Australia · galen.cuthbertson@anu.edu.au

Large Language Models (‘LLMs’) now dominate debates about technology and textuality across the humanities. Much of this work takes what we call a representationalist approach, which assumes that a model’s outputs reflect properties of their training data and can be investigated separately from their conditions of production and use. This paper reports on a practical inquiry into the limits of, and alternatives to, this representationalist understanding, pursued through masking experiments with BERT models finetuned on historical corpora. Building on a recent theoretical account of LLMs as demonstrating the probabilistic nature of humanistic textual practices, we offer a speculative experiment that explores the potential of embracing rather than resolving computational indeterminacy, and ask what it would mean to approach LLM outputs as situated enactments of historical probabilities.

Representationalism and LLMs

Representationalism has been extensively questioned within the history and philosophy of science as a basis for understanding modelling. Various arguments instead maintain that scientific models do not mirror or reflect existing systems but participate in their articulation (Knuuttila 2011; de Oliveira 2021). Even accounts that retain a representationalist component emphasise that models do far more than represent their targets, including enabling interference, coordinating practices, and enacting phenomena (Morrison & Morgan 1999; Cartwright 1999; Giere 2004). Yet representationalism remains the default framework through which word embeddings and LLMs are understood, in technical and humanistic arguments alike (Mikolov et al. 2013; Rogers et al. 2020).

In machine learning research, this framework is most apparent in the ‘probing’ paradigm: the systematic attempt to determine what linguistic, semantic, or world knowledge is encoded by the ‘internal representations’ of artificial neural networks (Kantamneni et al. 2025). Researchers treat models as containers for semantic content that can be read off from activation patterns or geometric relations in vector space (Heinzerling & Inui 2024; Engels et al. 2024; Li et al. 2025). This framing underpins both celebratory accounts in which LLMs “learn” rich and “highly interpretable” representations of syntax, semantic, and world knowledge (e.g. Cunningham et al. 2023) and critical accounts in which they fail to “ground” their representations in genuine meaning (e.g. Bender & Koller 2020; c.f. Mollo & Millière 2023). Representationalism similarly organises humanistic critiques of LLMs – as repositories of cultural or linguistic bias (Bonil et al. 2025) or as distorting or losing detail (Delétang et al. 2023) or meaning (Bender, Gebru, et al. 2021) in what they copy – and descriptions of them as powerful semantic search devices capable of surfacing latent cultural affinities and textual connections (Wilson et al. 2023).

This latter claim is made by the largest digital humanities project, to date, to pursue historical research with LLMs. The Living with Machines (‘LwM’) project, conducted at the British Library’s Turing Institute from 2018-2023, based many of its arguments on a BERT-derived model fine-tuned on nineteenth-century British Library books and journals (‘BLERT’). They used this model for various masked-token querying experiments, hiding a keyword and examining the model’s “unexpected” or “unusual” completions as evidence of usages in the underlying archive, and presenting changes in completions between models trained on different historical corpora as signalling shifts in cultural understandings and usages (Wilson et al. 2023). While the LwM arguments are careful with their findings, they generally assume that the model reproduces the latent patterns in the corpus closely enough that its completions can be taken as evidence of features of that corpus.

The limits of a representational experiment with BERT

Our experiments with the LwM approach demonstrate the practical limits of this representational framework. To explore Irishness in Australian literature, we began by fine-tuning the LwM’s BLERT model with text from To Be Continued (TBC) (Bode and Hetherington 2018-): a database of textual and bibliographical data for over 51,000 publications of novels, novellas and short stories published in historical Australian newspapers. We created an overall BLEAT (British Library-Australian BERT) model based on 100 years of TBC data (1840-1940), as well as five successive models trained on 20-year subsets of the data (1840-1859, 1860-1879, 1880-1899, 1900-1919 and 1920-1939).

For an initial masking experiment, we used 90 sentences from our corpus containing “Irish” or “Irishman” and examined the most likely completions from our different BLEAT models. The results included some correct predictions (14% overall) and, alongside some mundane responses (e.g. “man,” “men,” different nationalities), a tendency towards positive cultural associations (e.g. “musical,” “noble,” “humble,” “simple,” “sweet”). But when we turned to interpreting these individual results and what they suggested about the corpus or the language model, we found ourselves assuming causal mechanisms for what were probabilistic situations - an assumption undermined by the evident influence of the unmasked words in the query sentences on the results produced. While the LwM authors observe the influence of model opacity on their findings, in using specific results to support claims about cultural meaning and historical change they assume a much more direct and deterministic relationship between inputs and outputs of the model than we found reasonable.

Alternatives to representationalism

A range of theoretical arguments challenge representationalism without offering alternative approaches to employing LLMs for historical and textual research. Wendy Hui Kyong Chun’s Discriminating Data (2021) argues that algorithmic systems perpetuate social categories not by reproducing what training data “contains” but through operations that structure learning operations, including homophily, clustering, and optimisation. Phan and Wark (2021) make a parallel argument for race in Large Image Models, describing it not as a preexisting category reflected (or mis/represented) by the data, but as an epiphenomenon of algorithmic processes of classification, proxy selection, and inductive inference. Such arguments resonate with earlier digital humanities critiques of the limits of representationalism, in studies that explore how computational architectures and modelling procedures participate in the distinctions and relations through which race (McPherson 2012) and gender (Mandell 2019) become conceivable. The problem, for our purposes, is that while these accounts reconceptualise what models do in non-representational terms, they offer little practical guidance on how such insights can be translated into methods for working with LLMs.

This lack of methodological guidance is highlighted by a recent article by Fabian Offert and Ranjodh Singh Dhaliwal (2024) that highlights the representationalist assumptions structuring current critical AI methodologies through three “casuistries.” The “benchmark casuistry” connects most explicitly to the approach that we initially took with our results, and that is suggested by the LwM project: individual completions invite interpretations as textual evidence - a word surfaced from the archive or library collection, citable and traceable - when they were artefacts of a probabilistic process with no such evidentiary pathway. The “black box casuistry” reinforces this error, in treating LLMs as fixed objects with inner (opaque) workings, and the “stack casuistry” describes how this logic connects to assumptions about training data, in assuming a traceable pathway from input through to algorithmic operations and output. While the authors argue that LLMs demand different kinds of methodological engagement, their “propaedeutic” essay does not offer them. It does, however, point to an insight from Johanna Drucker that we have developed into an experiment with our BLEAT models.

An experiment in postrepresentational historical LLM research

Proposing what it might mean to operate at the “interface” of qualitative and quantitative critique,” Offert and Dhaliwal (2024: 4) refer to Drucker’s (2018) proposal “to understand the act of interpretation as the collapse of the probability space of all possible interpretations – to always keep in mind, in other words, that even the best anecdote is a sample.” This suggested a different approach to our BLEAT models: rather than selecting individual high-probability completions as findings, we extract full probability distributions over all possible substitutes, treating the entire vocabulary as the object of analysis. For each of our five temporal models, we analyse:

(1) mean probability and variance assigned to “Irish” itself,

(2) the distribution of high-probability alternatives; and

(3) divergence from the BLERT and BLEAT baseline models.

In thereby tracking which substitute tokens show consistent rises or falls in probability across each period, we examine the entire vocabulary as a probability distribution. Computing the Jensen-Shannon divergence between each period-specific model’s predicted distribution and that of the overall baseline gives us a coarse-grained measure of the period’s drift. While this single-scalar measure is a simple one, it offers a useful first orientation, indicating which periods diverge most sharply from the aggregate and where token-level inspection is likely to be most productive. We explore the results of this experiment, not as revealing the meaning of Irishness in different periods, but as engagements with the probabilistic, contextual, and apparatus-specific nature of LLM outputs.

Even with millions of probability assignments, no accumulation of completions will resolve the underlying epistemic and evidentiary difficulty. The problem is not that the models are small or the data is limited, but that each output is a completion within an apparatus that contingently configures text, query, and model, in ways that resist the forms of citation, corroboration, and archival location through which textual evidence is ordinarily incorporated into humanities arguments. Approaching LLM outputs in this way shifts attention from what a model “contains” to what practices of querying, fine-tuning, and interpreting bring forth. We conclude by suggesting that grappling with LLMs for historical inquiry may require abandoning the separation of “cultural object” and “humanistic interpretation,” and treating LLM outputs as situated performances of historical probabilities.

References
  1. Bender, Emily M. / Gebru, Timnit / McMillan-Major, Angelina / Shmitchell, Shmargaret (2021): “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜”, in: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, ACM: 610–623. DOI: 10.1145/3442188.3445922.
  2. Bender, Emily M. / Koller, Alexander (2020): “Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data”, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics: 5185–5198. DOI: 10.18653/v1/2020.acl-main.463.
  3. Bode, Katherine / Hetherington, Carol (eds.) (2018–): To Be Continued: The Australian Newspaper Fiction Database https://readallaboutit.com.au/.
  4. Bonil, Gustavo / Hashiguti, Simone / Silva, Jhessica / Gondim, João / Maia, Helena / Silva, Nádia / Pedrini, Helio / Avila, Sandra (2025): “Yet Another Algorithmic Bias: A Discursive Analysis of Large Language Models Reinforcing Dominant Discourses on Gender and Race.” Preprint, arXiv. DOI: 10.48550/ARXIV.2508.10304.
  5. Cartwright, Nancy (1999): The Dappled World: A Study of the Boundaries of Science. Cambridge: Cambridge University Press. DOI: 10.1017/CBO9781139167093.
  6. Chun, Wendy Hui Kyong (2021): Discriminating Data: Correlation, Neighborhoods, and the New Politics of Recognition. Cambridge, MA: The MIT Press.
  7. Cunningham, Hoagy / Ewart, Aidan / Riggs, Logan / Huben, Robert / Sharkey, Lee (2023): “Sparse Autoencoders Find Highly Interpretable Features in Language Models.” Preprint, arXiv. DOI: 10.48550/ARXIV.2309.08600.
  8. de Oliveira, Guilherme Sanches (2021): “Representationalism Is a Dead End”, in: Synthese 198, 1: 209–235. DOI: 10.1007/s11229-018-01995-9.
  9. Delétang, Grégoire / Ruoss, Anian / Duquenne, Paul-Ambroise / Catt, Elliot / Genewein, Tim / Mattern, Christopher / Grau-Moya, Jordi / Li, Li Kevin / Kalle Kossaifi, Jean / Gurnee, Wes / Legg, Shane (2023): “Language Modeling Is Compression.” Preprint, arXiv. DOI: 10.48550/ARXIV.2309.10668.
  10. Drucker, Johanna (2018): The General Theory of Social Relativity. The Elephants Ltd.
  11. Engels, Joshua / Michaud, Eric J. / Liao, Isaac / Gurnee, Wes / Tegmark, Max (2024): “Not All Language Model Features Are One-Dimensionally Linear.” Preprint, arXiv. DOI: 10.48550/ARXIV.2405.14860.
  12. Engels, Joshua / Riggs, Logan / Tegmark, Max (2024): “Decomposing the Dark Matter of Sparse Autoencoders.” Preprint, arXiv. DOI: 10.48550/ARXIV.2410.14670.
  13. Giere, Ronald N. (2004): “How Models Are Used to Represent Reality”, in: Philosophy of Science 71, 5: 742–752. DOI: 10.1086/425063.
  14. Heinzerling, Benjamin / Inui, Kentaro (2024): “Monotonic Representation of Numeric Properties in Language Models.” Preprint, arXiv. DOI: 10.48550/ARXIV.2403.10381.
  15. Kantamneni, Subhash / Engels, Joshua / Rajamanoharan, Senthooran / Tegmark, Max / Nanda, Neel (2025): “Are Sparse Autoencoders Useful? A Case Study in Sparse Probing.” Preprint, arXiv. DOI: 10.48550/ARXIV.2502.16681.
  16. Knuuttila, Tarja (2011): “Modelling and Representing: An Artefactual Approach to Model-Based Representation”, in: Studies in History and Philosophy of Science Part A 42, 2: 262–271. DOI: 10.1016/j.shpsa.2010.11.034.
  17. Li, Yuxiao / Michaud, Eric J. / Baek, David D. / Engels, Joshua / Sun, Xiaoqing / Tegmark, Max (2025): “The Geometry of Concepts: Sparse Autoencoder Feature Structure”, in: Entropy 27, 4: 344. DOI: 10.3390/e27040344.
  18. Mandell, Laura (2019): “Gender and Cultural Analytics: Finding or Making Stereotypes?”, in: Gold, Matthew K. / Klein, Lauren F. (eds.): Debates in the Digital Humanities 2019. Minneapolis: University of Minnesota Press. DOI: 10.5749/j.ctvg251hk.
  19. McPherson, Tara (2012): “Why Are the Digital Humanities So White? Or Thinking the Histories of Race and Computation”, in: Gold, Matthew K. (ed.): Debates in the Digital Humanities. Minneapolis: University of Minnesota Press. DOI: 10.5749/minnesota/9780816677948.003.0017.
  20. Mikolov, Tomas / Chen, Kai / Corrado, Greg / Dean, Jeffrey (2013): “Efficient Estimation of Word Representations in Vector Space.” Preprint, arXiv. DOI: 10.48550/ARXIV.1301.3781.
  21. Mollo, Dimitri Coelho / Millière, Raphaël (2023): “The Vector Grounding Problem.” Preprint, arXiv. DOI: 10.48550/ARXIV.2304.01481.
  22. Morrison, Margaret / Morgan, Mary S. (1999): “Models as Mediating Instruments”, in: Morgan, Mary S. / Morrison, Margaret (eds.): Models as Mediators. 1st ed. Cambridge: Cambridge University Press. DOI: 10.1017/CBO9780511660108.003.
  23. Offert, Fabian / Dhaliwal, Ranjodh Singh (2024): “The Method of Critical AI Studies, a Propaedeutic.” Preprint, arXiv. DOI: 10.48550/ARXIV.2411.18833.
  24. Phan, Thao / Wark, Scott (2021): “Racial Formations as Data Formations”, in: Big Data & Society 8, 2: 20539517211046377. DOI: 10.1177/20539517211046377.
  25. Rogers, Anna / Kovaleva, Olga / Rumshisky, Anna (2020): “A Primer in BERTology: What We Know About How BERT Works”, in: Transactions of the Association for Computational Linguistics 8: 842–866. DOI: 10.1162/tacl_a_00349.
  26. Wilson, Daniel CS / Coll Ardanuy, Mariona / Beelen, Kaspar / McGillivray, Barbara / Ahnert, Ruth (2023): “The Living Machine: A Computational Approach to the Nineteenth-Century Language of Technology”, in: Technology and Culture 64, 3: 875–902. DOI: 10.1353/tech.2023.a901591.