Daejeon, July 27–31
Human memory is fundamentally an apparatus for mapping the world and orienting ourselves within it. Cognitive scientists have characterized the brain as a predictive machine, constantly using stored knowledge to anticipate and interpret incoming information (Clark, 2023). This mapping function of memory is also collective. Cultural and social memories serve to orient groups within the larger human history. Human societies create collective memories through external symbols and records (monuments, texts, archives). This cultural memory transcends individual lifespans by storing knowledge in durable forms that outlast communicative memory passed by word-of-mouth (Halbwachs, 1925; Assmann, 2008). Such externalized memory allows each generation to inherit a mapped landscape of historical experience: it structures how communities engage with the past, which stories are made accessible, and which experiences remain difficult to reach.
Our paper extends this concept to large AI models, trained on vast archives of human knowledge and culture and increasingly used to mediate access to cultural records. These models assimilate cultural and historical data across centuries and make it accessible through interactive prompts, rather than through conventional search or retrieval alone. We approach them as cognitive memory agents: systems whose operations of retrieval, synthesis, and generation reshape the conditions under which the past becomes accessible (Schuh, 2024). They act as compasses within cultural memory, orienting users through paths, associations, summaries, and hierarchies of relevance shaped by training data, model design, and platform governance. Building on cultural memory studies, digital memory studies, media history, cultural analytics, web historiography, platform capitalism, and AI-mediated remembering, we extend the analysis of changing media environments to AI systems understood as new infrastructures of orientation. These infrastructures are never neutral: they shape what can be found, how it is framed, whose records are amplified, and which forms of knowledge remain marginal or opaque. We therefore propose “artificial memory” as a conceptual framework for understanding AI-mediated access to cultural records, while identifying concrete points where DH can intervene: datasets, retrieval workflows, interface design, provenance mechanisms, and governance.
The challenge of navigating ever-expanding knowledge is not new. Throughout history, humans have developed external memory aids and organizational systems to orient themselves in an overwhelming memory of recorded information. French anthropologist André Leroi-Gourhan chronicled the progressive externalization of memory from oral traditions to writing to modern data storage (Leroi-Gourhan, 2018). This argument belongs to a longer genealogy of media understood as extensions or prostheses of human capacities, from Plato’s ambivalent account of writing in the Phaedrus to McLuhan’s theory of media as extensions of the body and Stiegler’s account of technical memory (Plato, 1995; McLuhan, 1964; Stiegler, 1998).
As the archive of human knowledge expanded, memory had to be augmented by external tools (Kamin, 2021). The digital revolution accelerated this externalization process (Hoskins, 2017). Search engines became a form of universal index allowing instant retrieval of facts and sources (Vaidhyanathan, 2011, Renouard, 2017). The move from catalogues to search engines transformed how scholars and the public engage with cultural memory, moving from guided, institutionally framed pathways to more individualized, query-driven encounters (Putnam, 2016). Their main function, however, remained to direct users toward documents, records, or ranked results.
Generative AI changed the user’s relation to those traces. Earlier search systems were already dynamic and probabilistic, but they generally returned sources to be explored. Large Language Models produce conversational syntheses instead, drawing on probabilistic relationships learned from large collections of books, articles, web pages, and cultural records. These relationships form a latent map of cultural knowledge, an implicit atlas where concepts and facts are embedded as vectors and clustered by semantic similarity and contextual usage (Underwood, 2021). Pre-trained language models can therefore operate as massive knowledge bases, with the advantage of flexible querying in natural language across an open domain of topics (Petroni et al., 2018). Yet this form of mediation can make the path from source to answer difficult to reconstruct. The AI becomes an active memory agent within our collective knowledge ecosystem.
While archivists and librarians have long worked actively to preserve, organize, and animate collections, their core mandate has historically centered on safeguarding materials and providing mediated access through catalogues, finding aids, and reading-room practices. With AI systems now being trained on large-scale digitized collections, we are seeing an additional layer emerge in which archives increasingly function as computational knowledge services. Michael Moss et al. (2018) note that the digital abundance of records pushes archivists and users to see cultural objects “as collections of data to be mined and not as texts to be read” (Moss et al., 2018). Archives in the digital age are curated by humans, but also queried, sampled, and modeled by algorithms. For historians and archivists, this raises practical and epistemic challenges: the volume of digital records transforms historical research, forcing historians to rely on search engines, databases, and increasingly automated tools in order to cope with web-scale corpora (Milligan, 2019 and 2022). These tools, including AI-based systems, alter how sources are discovered, read, and interpreted.
Using AI systems as memory tools reflexively transforms our own cognitive practices. When scholars, students, or the public use LLM-based assistants, they begin to offload certain memory tasks to these “external brain”; what some theorists call exosomatic memory or exograms (Welzer, 2008). Yet the difference is that whereas a book or a database sits idle until a person consults it, an AI agent proactively shapes the knowledge retrieval process. It selects, summarizes, and even “decides” which aspects of memory to present.
Carl Öhman even suggests that AI gives the past a form of agency (Öhman, 2025). Digital humanities scholars have begun to explore this, for example by creating AI chatbots that impersonate historical figures, offering new forms of interaction between audiences and historical materials (Natale et al., 2025), or by employing machine learning to detect themes and biases in historical newspapers.Carl Öhman even suggests that AI gives the past a form of agency (Öhman, 2025). Digital humanities scholars have begun to explore this, for example by creating AI chatbots that impersonate historical figures, offering new forms of interaction between audiences and historical materials (Natale et al., 2025), or by employing machine learning to detect themes and biases in historical newspapers.
But the question is not whether archives were previously “passive” and are now “active,” but rather how new computational layers redistribute agency among archivists, users, and algorithms in the ongoing production of collective memory.
One major question of AI-mediated memory is that of who controls the configuration and distribution of collective memory. Whereas archives and libraries are typically governed by public institutions or scholarly principles (at least ideally, with missions of preservation, access, neutrality), the most advanced AI models today are often developed by private corporations with proprietary interests. Tech companies are already positioning these AI platforms at the center of information-seeking, integrating them into search engines, virtual assistants, and educational tools. In such contexts, the first encounter with a cultural object may no longer be a catalogue entry, archival description, or digitized source, but an AI-generated synthesis produced within a proprietary interface. This platformization of memory raises concerns about bias, transparency, and accountability (Smit, 2025; Kasy, 2025; Durand Folco and Martineau, 2023).
Large models are often trained on digitized cultural heritage scraped en masse from the internet or databases. This continues earlier tensions around mass digitization, where cultural access, infrastructural dependency, and corporate control were already entangled (Thylstrup, 2018). If AI models are built on such data without permission or context, it can be seen as a new form of appropriation. Moreover, if the resulting model is proprietary, the flow of knowledge goes one way: heritage is harvested to create a commercial AI that may not freely give back what it has learned. This asymmetry can be read as part of a broader logic of data colonialism, in which cultural and social traces are extracted, reorganized and monetized by platforms that do not necessarily remain accountable to the communities from which those traces originate (Couldry and Mejias, 2019).Large models are often trained on digitized cultural heritage scraped en masse from the internet or databases. This continues earlier tensions around mass digitization, where cultural access, infrastructural dependency, and corporate control were already entangled (Thylstrup, 2018). If AI models are built on such data without permission or context, it can be seen as a new form of appropriation. Moreover, if the resulting model is proprietary, the flow of knowledge goes one way: heritage is harvested to create a commercial AI that may not freely give back what it has learned. This asymmetry can be read as part of a broader logic of data colonialism, in which cultural and social traces are extracted, reorganized and monetized by platforms that do not necessarily remain accountable to the communities from which those traces originate (Couldry and Mejias, 2019).
Another challenge is the risk of hegemonic memory. AI systems tend to amplify dominant patterns present in their training data (Smit, Smits and Merrill, 2024). This risk is especially acute when minoritized languages, colonial archives, local collections, oral histories and community-held records are absent, poorly described, or absorbed into models without appropriate provenance and consent (Carroll et al., 2020). This concern echoes warnings that working with large cultural datasets can give a false sense of completeness (Manovich, 2020; Gupta and Kapoor, 2020; Lenardo and Kaplan, 2023). While the challenges continue to grow, DH offers opportunities to intervene.
There is a call within digital humanities and archival circles to govern these new memory devices more democratically: to demand transparency in training data, to require models to acknowledge sources, or even to develop public and open models that serve as commons-based memory rather than corporate-controlled memory (Hess and Ostrom, 2006). Models, datasets and interfaces are sites where values are negotiated (Jo and Gebru, 2020), and DH can help ensure that generative systems support rather than undermine diverse communities’ claims over their own memories. This could involve documenting dataset provenance, identifying representational gaps, allowing communities to define access conditions for sensitive materials, and evaluating model outputs with the participation of the groups whose histories are being represented.
Some approaches, like Retrieval-Augmented Generation (RAG), attempt to reattach generative models to verifiable sources by coupling them with authenticated databases and citation mechanisms (Lewis et al., 2020). Such an interface can display the generated answer alongside ranked source passages, collection metadata, confidence indicators, provenance warnings, and links back to the archival object or catalogue record. It can also separate what the model infers from what the source explicitly states, making the epistemic status of each claim visible to the user. A DH-oriented RAG workflow might begin with a curated corpus from a public archive or museum collection, enrich it with stable identifiers and metadata, index both texts and images as retrievable units, and then require the generative model to ground its response in those units rather than in an undifferentiated training memory.
But even with such techniques, the governance of digital collective memory must still confront a central question: on what grounds can we trust an AI system to represent our cultural record faithfully enough to use it as a memory device? Any answer will require new frameworks of verification, mechanisms of community oversight, and perhaps a redefinition of what counts as “authoritative” memory. Memory in the age of AI is likely to become probabilistic and negotiated rather than the traditional archival ideal of fixed, singular evidence.
In this shifting landscape, digital humanities can help recalibrate the compass. Rather than letting commercial platforms silently define the latent maps through which we navigate the past, DH scholars, archivists, and curators can intervene at several levels: by insisting on open, well-documented datasets; by building their own systems on top of public archives and museum collections; by designing interfaces that foreground provenance and epistemic status instead of hiding them; and by deepening collaborations between research infrastructures and heritage institutions. It involves building sustained partnerships between researchers, memory institutions, and affected communities to decide together how artificial memory systems should point. In this sense, DH can act as a compass-maker: designing instruments of orientation while also testing where existing systems are skewed by ownership, opacity, uneven provenance, and hegemonic training data.In this shifting landscape, digital humanities can help recalibrate the compass. Rather than letting commercial platforms silently define the latent maps through which we navigate the past, DH scholars, archivists, and curators can intervene at several levels: by insisting on open, well-documented datasets; by building their own systems on top of public archives and museum collections; by designing interfaces that foreground provenance and epistemic status instead of hiding them; and by deepening collaborations between research infrastructures and heritage institutions. It involves building sustained partnerships between researchers, memory institutions, and affected communities to decide together how artificial memory systems should point. In this sense, DH can act as a compass-maker: designing instruments of orientation while also testing where existing systems are skewed by ownership, opacity, uneven provenance, and hegemonic training data.