DH 2026

Daejeon, July 27–31

Wed, July 2909:00–10:30S098204-205
Long Paper

AI as Archive: Memory, Curation, and Reinterpretation in Large Language Models

Geethapriya Chandramohan
Indian Institute of Technology Madras, India · hs22d020@smail.iitm.ac.in

Introduction: The Archive After the Linguistic Turn in AI

Large language models have quietly become part of the infrastructure through which knowledge about the past is produced. They are not search engines, and not databases. They respond to questions by constructing plausible accounts, often with the tone and authority of archival knowledge. That shift has significant implications for how memory operates in digital environments.

This paper argues that LLMs should not be understood as neutral repositories of information. They participate actively in the production, curation, and suppression of cultural memory. Their architectures, training practices, and generative outputs raise fundamental questions about what counts as an archive, who controls it, and whose knowledge is preserved.

These questions have been building in digital humanities for some time. Roopika Risam (2019) has shown how digital archives reproduce colonial hierarchies through selective digitisation and the metadata standards of Western institutions. Nishant Shah (2019) documents how local memory formations in the Global South get marginalised by globally dominant data infrastructures. Critical algorithm studies, Noble (2018), Gillespie (2014), have tracked how the technical decisions embedded in digital systems carry social consequences that their designers often do not acknowledge. LLMs push all of this further, because they do not just store knowledge: they generate it.

Methodologically, the paper undertakes a conceptual analysis of publicly documented LLM architectures and training pipelines, focusing on how data selection, probabilistic generation, and platform governance reshape archival logic. Rather than offering a technical audit, it establishes a critical interface between computational processes and archival theory.

LLMs as Meta-Archives: Connective Memory and Generative Curation

Andrew Hoskins (2011) argues that contemporary digital memory is no longer a stable repository but an emergent process shaped continuously by algorithmic mediation.LLMs represent a particularly clear instance of this shift. Their memory is not stored content, it is pattern. What gets produced in response to a prompt is not retrieval but recomposition: something that feels like memory without being one.

This is what makes them meta-archives rather than archives in any conventional sense. Models like GPT-4 and LLaMA 2 are trained on web-scale corpora, Common Crawl, Wikipedia, digitised books, that represent a historically contingent and globally uneven sample of human knowledge production. Ask such a model about the 1947 Partition of India, and it does not pull up a document. It generates a statistically probable reconstruction from whatever patterns were present in its training data. The archive becomes non-locatable and unreproducible. Each response is, in a real sense, the only instance of itself.

N. Katherine Hayles (2017) argues that cognition circulates across networks of biological and computational agents rather than residing in any single mind or machine, a framework that maps usefully onto LLM outputs, which are shaped not just by internal algorithms but by user prompts, training data provenance, infrastructure constraints, and platform governance decisions. A key distinction must be maintained here: some accounts collapse Hayles’s framework with Braidotti’s posthumanism as though they articulate the same position. They are not. Hayles foregrounds adaptive cognition and distributed agency; Braidotti is working from an affirmative ethics and a relational ontology. The distinction matters for how we assign responsibility when these systems go wrong.

Posthumanism, Bias, and the Politics of Archival Memory

Braidotti (2013) offers tools for thinking about LLMs as posthuman formations, systems where memory, creativity, and interpretation have outgrown their exclusively human frame. That is genuinely interesting. While this is theoretically generative, it risks aestheticising dynamics that require closer scrutiny: these models are trained on corpora saturated with the residues of centuries of patriarchy, colonialism, casteism, and epistemic exclusion. The posthuman archive does not escape history. It compresses it, and smooths over the texture in the process.

The evidence is not abstract. Studies have found that GPT-3 consistently associates African American Vernacular English with negative sentiment and defaults to male pronouns in professional role completions. Bender et al.’s (2021) critique of LLMs as “stochastic parrots” is especially relevant here: their analysis of the structural risks of large-scale models, and the conditions under which that work was suppressed at Google, constitutes a case study in who gets to determine what AI systems know and remember. Trained overwhelmingly on English-language web data, these systems flatten global linguistic diversity. Gebru et al.’s (2021) parallel work on dataset documentation underscores the same problem from a transparency standpoint: without systematic accountability for what data enters training pipelines, the exclusions remain invisible. Tamil, Swahili, and Quechua remain severely underrepresented. What is lost is not only vocabulary but the cultural memory those languages carry, with no indication in the output to signal this absence.

Postcolonial and Dalit scholarship makes this concrete. Contributors to the Digital Humanities Quarterly special issue on postcolonial digital humanities have documented how caste hierarchies are reproduced in the metadata structures of Indian digitisation projects and in the near-absence of Dalit-authored texts from training corpora. Risam’s (2019) framework bears repeating: the archive has never been neutral, and shifting from institutional to algorithmic curation does not fix that. In some respects, it makes the problem harder to see, because the bias is now embedded in weights rather than policies.

Infrastructure, Classification, and Archival Invisibility

Bernard Stiegler (1998) traces a long process by which memory has been progressively externalised, into writing, then photography, then digital storage. LLMs represent a qualitative break in that sequence. In earlier forms, the memory support and the interpretive act were at least separable. With LLMs, they collapse into the same operation. Every output is simultaneously a storage act and an interpretive one, and these cannot be separated after the fact.

Bowker and Star’s (1999) work on classification is directly applicable. The datasets used to train LLMs are assembled through opaque processes of scraping, filtering, tokenising, and weighting. Those choices determine what knowledge surfaces in model outputs. The technical framing makes it look like necessity, like the data is simply what was available. This is not the case. It is a set of political decisions that have been absorbed into infrastructure and naturalised as if they were neutral or inevitable. Gillespie (2014) makes a parallel point about content moderation: the decisions are rarely disclosed, but they shape everything downstream.

Wendy Hui Kyong Chun’s (2008) concept of enduring ephemerality adds a temporal dimension. LLMs present an illusion of stable knowledge while relying on continuous updates and retraining. The transition from GPT-3.5 to GPT-4 changed what the system effectively “knew” and how it behaved, with no stable public record of what changed or what may have been lost. Memory here is not preserved, it is perpetually rewritten, without an audit trail. The implications for cultural inheritance and archival accountability are serious, and they remain largely unaddressed.

Archive Fever and the Generative Paradox

Derrida (1996) describes the archive as fundamentally contradictory: driven simultaneously by the desire to preserve and the compulsion to reorder. LLMs may be the most literal instantiation of that paradox so far, because the two impulses cannot be separated. Every output is recall and transformation at once. There is no retrieving a stable record, only producing a new one, each time as a new instantiation.

For DH practitioners, the consequences are concrete. Projects like the Europeana network or the British Library’s Living with Machines initiative that use LLMs for metadata generation or cross-lingual summarisation are not simply describing archival content, they are actively re-curating it, whether they frame it that way or not. Scholars who cite LLM-generated outputs face a provenance problem that existing citation standards are not built to handle: what the model generated cannot be retrieved, reproduced, or independently verified the way archival records can. This is not a minor methodological inconvenience. It is a structural problem.

Implications for Digital Humanities Practice

Understanding LLMs as archival assemblages requires a rethinking of DH practice in three key areas.

First, archival design must critically assess the use of LLMs in curation and description. These processes should be transparent, documented, and open to scrutiny, rather than being treated as neutral automation. The opacity currently surrounding these systems is a policy choice, not a technical constraint.

Second, knowledge production in Global South and postcolonial contexts demands particular attention. Scholars working on multilingual archives, oral histories, and community memory projects must resist the homogenising effects of LLM-assisted curation by insisting on community-controlled metadata standards and genuine multilingual data sovereignty.

Third, scholarly standards must adapt to LLM-generated content. Clear distinctions are needed between retrieved archival material and generated outputs, along with new norms for citation and verification. This distinction is therefore essential.

Conclusion

Large language models are not simply tools that interact with archives; they are archival formations that reshape the conditions under which memory is produced, accessed, and interpreted. By foregrounding generativity over storage, opacity over custodianship, and recombination over retrieval, they transform the archive into a dynamic and contested space.

The stakes of this transformation are ethical and political. As LLMs increasingly mediate access to knowledge, the question of whose histories they preserve or erase becomes urgent. Posthumanist, feminist, decolonial, and Dalit frameworks are not methodological extras here, they are the instruments best suited to seeing what is actually happening inside these systems, and to imagining the kind of archives that might do otherwise.

In this sense, the archive in the age of AI is no longer a repository of the past but a site of ongoing negotiation: one in which cultural memory is continuously reconfigured through the interplay of human and machinic processes.

References
  1. References
  2. Bender, Emily M. / Gebru, Timnit / McMillan-Major, Angelina / Shmitchell, Shmargaret (2021): On the dangers of stochastic parrots: Can language models be too big?, in: Proceedings of FAccT 2021: 610–623.
  3. Bowker, Geoffrey C. / Star, Susan Leigh (1999): Sorting Things Out: Classification and Its Consequences. Cambridge, MA: MIT Press.
  4. Braidotti, Rosi (2013): The Posthuman. Cambridge: Polity Press.
  5. Burgess, Jean / Green, Joshua (2018): YouTube: Online Video and Participatory Culture. 2nd ed. Cambridge: Polity Press.
  6. Chun, Wendy Hui Kyong (2008): The enduring ephemeral, or the future is a memory, in: Critical Inquiry 35, 1: 148–171.
  7. Derrida, Jacques (1996): Archive Fever: A Freudian Impression. Trans. E. Prenowitz. Chicago: University of Chicago Press.
  8. Gebru, Timnit / Morgenstern, Jamie / Vecchione, Briana / Wortman Vaughan, Jennifer / Wallach, Hanna / Daumé, Hal / Crawford, Kate (2021): Datasheets for datasets, in: Communications of the ACM 64, 12: 86–92.
  9. Gillespie, Tarleton (2014): The relevance of algorithms, in: Gillespie, Tarleton / Boczkowski, Pablo J. / Foot, Kirsten A. (eds.): Media Technologies: Essays on Communication, Materiality, and Society. Cambridge, MA: MIT Press 167–194.
  10. Hayles, N. Katherine (2017): Unthought: The Power of the Cognitive Nonconscious. Chicago: University of Chicago Press.
  11. Hoskins, Andrew (2011): Media, memory, metaphor: Remembering and the connective turn, in: Parallax 17, 4: 19–31.
  12. Noble, Safiya Umoja (2018): Algorithms of Oppression: How Search Engines Reinforce Racism. New York: New York University Press.
  13. Risam, Roopika (2019): New Digital Worlds: Postcolonial Digital Humanities in Theory, Praxis, and Pedagogy. Evanston, IL: Northwestern University Press.
  14. Shah, Nishant (2019): Digital Humanities on the Ground: Post-Access Politics and the Second Wave of Digital Humanities, in: South Asian Review 40, 3: 155–173. DOI: 10.1080/02759527.2019.1599551.
  15. Stiegler, Bernard (1998): Technics and Time, 1: The Fault of Epimetheus. Trans. R. Beardsworth and G. Collins. Stanford: Stanford University Press.