Daejeon, July 27–31
Digital archives in the humanities have traditionally been designed around a static information retrieval (IR) paradigm, in which users submit predefined queries and receive fixed sets of results. Classical information retrieval research has already identified the limitations of this model, noting that real-world information seeking is iterative, exploratory, and context-dependent rather than linear and query-bound (Bates 1989).
Within Digital Humanities and archival studies, this critique has led to calls for alternatives to search-centric archival interfaces. Archival theory has challenged the notion of archives as neutral, passive repositories, emphasizing instead their role as active sites where memory, power, and interpretation are continuously negotiated (Schwartz and Cook 2002). In parallel, Whitelaw (2015) argues that conventional digital collections privilege keyword search at the expense of contextual exploration, proposing “generous interfaces” that foreground discovery and relational understanding.
Building on these critiques, researchers have proposed models such as Participatory Archives, which reconceptualize users as active contributors to archival meaning-making rather than passive consumers (Huvila 2008), and Living Archives, which frame archives as dynamic environments that connect historical materials with present-oriented creative and performative practices (Sabiescu 2020; Almeida and Hoyer 2020). However, despite their strong theoretical grounding, these approaches have often remained only partially realized at the level of actual user interaction and system design.
This study aims to examine how the integration of Large Language Models (LLMs) into a digital literary archive can operationalize and extend the theoretical shift from static information retrieval toward participatory and living archival practices. Specifically, it investigates whether LLM-based interfaces enable forms of engagement that substantively realize the ideals of participatory and living archives, rather than merely rearticulating them at a conceptual level.
The study analyzes user interaction logs from Humanitext Aozora, a Retrieval-Augmented Generation (RAG) system built on Aozora Bunko, a major Japanese digital library of public-domain literature (https://aozora.humanitext.ai/). The raw dataset consisted of 63,392 user interactions collected over a four-month period. After excluding custom-mode prompts and entries from a legacy interaction schema (3,219 records; 5.08% of the data), whose heterogeneous structure made consistent intent classification difficult, a total of 60,173 interactions were retained for analysis. These queries were examined through a mixed-methods approach combining quantitative classification and qualitative analysis in order to identify recurring interaction patterns and underlying user intentions.
The analysis reveals that user engagement with the LLM-integrated archive extends far beyond conventional information retrieval. User interactions can be classified into three dominant modalities:
(1) Informational Retrieval (24.66%; 14,837 queries), encompassing requests for factual information, bibliographic data, summaries, and explanatory question-and-answer interactions related to literary works.
Examples:
“Please list ten literary works that engage with the philosophy of Arthur Schopenhauer.”
“What Freud-inspired motifs can be identified in the works of Shinichi Makino?”
(2) Persona Simulation (40.03%; 24,090 queries), in which users engage in dialogic interactions with simulated literary authors or characters, addressing them as conversational partners within an imagined narrative or authorial persona.
Examples:
“You are Ryōtarō Shiba. What do you think about cryptocurrencies such as Bitcoin?”
“Let’s role-play as lovers. You take the male role. We will begin now.” (addressed to Osamu Dazai)
(3) Stylistic Emulation and Derivative Creation (35.31%; 21,246 queries), where users instruct the system to generate new texts by emulating the stylistic features of historical authors or by producing derivative narratives such as continuations, rewritings, or reinterpretations of existing works.
Examples:
“Please rewrite the opening of Kusamakura by Natsume Sōseki in a manner inspired by the concept of ‘absolute contradictory self-identity.’” (addressed to Kitarō Nishida)
“Imagine Rashōmon if its servant were a modern office worker.” (addressed to Ryūnosuke Akutagawa)
Notably, interactions oriented toward persona simulation and stylistic or derivative creation together constitute over 75% of all classified queries, indicating that users predominantly engage with the archive not as a search interface, but as a dialogic and generative literary environment. These findings suggest that the LLM does not merely enhance search efficiency, but rather transforms the archive into a generative and dialogic environment. In doing so, it concretely instantiates principles long articulated in participatory and living archive theories, enabling users to actively reinterpret, reenact, and recontextualize literary heritage.
We argue that such interactions constitute a distinct mode of literary engagement, which we term “generative reception.” In this mode, the archive functions not only as a source of evidence but as a substrate for experimentation, where the boundary between reading and rewriting becomes increasingly permeable. By empirically demonstrating how LLM integration reshapes user behavior, this study contributes to ongoing debates in Digital Humanities and archival studies concerning the future of archives as dynamic, participatory, and living systems. In this sense, generative reception does not merely represent a new user behavior, but signals a structural transformation of the archive itself—from an interface for retrieval to an environment for literary experimentation.