DH 2026

Daejeon, July 27–31

Wed, July 2916:30–18:00S077107
Short Paper

Augmented Reading. Comparing Human Behaviors and Large Language Models in Reading Experience

Gino Roncaglia
Università degli Studi "Roma Tre", Italy · gino.roncaglia@uniroma3.it
Ludovica Mastrobattista
Università degli Studi "Roma Tre", Italy · ludovica.mastrobattista@uniroma3.it
Elena Ranfa
Università degli Studi "Roma Tre", Italy · elena.ranfa@uniroma3.it
Lorenzo Picca
Università degli Studi "Roma Tre", Italy · lorenzo.picca@uniroma3.it
Francesco Leonetti
AI Consultant Specialist · leonetti.f@gmail.com

The pervasiveness of the digital ecosystem has radically transformed reading practices and habits (Baron 2015; 2021), generating new forms of integration between text consumption and online research aimed at gathering further information or vocabulary enrichment. This transformation aligns with the interpretative framework of reading as a dynamic and cooperative process between text and reader (Iser 1978; Eco 1979; Suleiman / Crosman 1980; Franzmann 1999), nowadays profoundly reshaped by the hypertextual, networked and radial dimensions of digital environments (McGann 2001; Landow 2006). Recent discussions on non-linear and networked writing and reading, including forms of hyperlink-driven textuality (Rettberg 2018; Tabbi 2018; Salter / Moulthrop 2021), further reinforce the view of reading as an expansive, multi-directional activity shaped by digital infrastructures. Within this perspective, “augmented reading” refers to the reader’s exploratory behaviors induced by trigger words or expressions, considered as possible exit points from the text to autonomous on-line research of external content (Roncaglia 2022). This concept is grounded in the open nature of texts and – reminiscent of the semiotic theory of intertextuality (Orr 2003; Mason 2019; Allen 2021) – distinguishes three types of references: intertextual, intermedial, and extratextual, thereby reevaluating pragmatic and non-linear reading behaviors that are often overlooked (Roncaglia 2020). Triggers may include names of people, places, bibliographic references, as well as less obvious elements that prompt readers to pause and seek additional information online. This approach emphasizes the active role of the reader in expanding text comprehension, both by following explicit references and through autonomous discoveries, reflecting the increasing complexity of reading paths and their networked nature (Faggiolani / Vivarelli 2016; Faggiolani / Verna / Vivarelli 2017). Such reading practices demonstrate how textual content consumption is structured through dynamics of connection, expansion, and semantic recomposition. In this perspective, systematic analysis of exit points from the text can provide valuable insights into reading behaviors, practices, and habits in the new digital ecosystem (Mastrobattista 2024), which can also inform the design of more conscious and personalized forms of publishing (Ranfa 2023).

The research developed within the CHANGES project

Project funded by PNRR funds – Mission 4, Component 2, Investment 1.3 – CHANGES (Cultural Heritage Active Innovation for Next-Gen Sustainable Society) – PE00000020 – Spoke 3 – Funded by the European Union – NextGenerationEU – CUP Code: F83C22001650006

investigates the potential of using Large Language Models (LLMs) as tools to analyze reading behaviors, considering artificial intelligence as a new, and potentially useful, model for interpreting our textual practices (Roncaglia 2023; 2025) and possibly as model reader (Ciotti 2024). Beyond its empirical scope, the study addresses a core question: whether LLMs can approximate interpretative selection processes traditionally associated with reader-response theory, thereby exploring the epistemic and hermeneutic implications of generative systems within digital humanities. More specifically, the aim of the research is to determine the extent to which the exit points identified by artificial intelligence models coincide – or significantly approximate – those identified by human readers during a reading and annotation task on four texts. The experiment draws on empirical data collected by the Italian PRIN (Research Projects of National Relevance) NEREIDE (New Reading Experiences in the Digital Ecosystem), which studied new reading habits and practices among different groups of human readers, enabling a comparison between human and artificial behaviors.

The human experiment analyzed the actions that accompany and follow reading and involved three groups of readers: secondary school students (15-18 years), university students (18-34 years), and adult readers (≥35 years). Participants read four excerpts of essays – two narrative and two argumentative – with different expected degrees of intertextuality (two with an expected high degree of intertextuality, two with an expected lower degree), applying specific annotation methods to interact with the text and indicate potential exit points. The annotation scheme distinguished between potential exit points, marking moments of attention without interrupting the reading flow, and actual exit points, marking a temporary interruption of reading. Additionally, the experiment employed eye-tracking technology to detect fixation points during on-screen reading of the same texts, following a methodology that allows observation of the complexity of cognitive processes associated with reading (Wolf 2007; 2018; Johns 2023). The results of the NEREIDE project show a strong correlation between the exit points explicitly marked by readers and the fixation points detected through the eye-tracking, as well as a strong correlation between the expected degree of intertextuality of the text used and the number of exit points detected.

Starting from those results, for the CHANGES project a research protocol was developed to replicate, using LLM, the conditions of the human task. The four texts used with human readers were submitted to three different Cloud-based LLMs (ChatGPT-5, Claude Sonnet, and Gemini 3 Advanced) through a structured multi-phase process. The meta-prompt used was carefully designed to replicate the instructions given to human participants, asking the models to identify and classify potential exit points according to the same categories used in the annotation task: references to people, places, historical events, bibliographic citations, specialized terms, and concepts likely to stimulate further research. To ensure comparability of results, each model was queried multiple times on the same text, with slight variations in prompt formulation to test response stability. This strategy aligns with recent findings showing how appropriately engineered LLMs can reliably support qualitative analysis tasks, even when deployed locally, thanks to advanced prompt-engineering techniques (Meyer et al. 2025). The outputs generated by the LLMs were then processed and normalized to enable quantitative and qualitative comparison with the empirical data collected in the first phase. This included analyzing the frequency of trigger identification, their type and distribution within the texts, as well as evaluating the semantic correspondence between the points identified by the models and those highlighted by human participants.

The comparison between the results produced by LLM and the empirical data from the human experiment aimed to assess whether similar exit points are identified by human readers and by artificial intelligence under comparable task conditions. A significant correspondence between the behaviors of Large Language Models and those of human readers would pave the way for the use of language models as innovative tools not just for analyzing and evaluating, but even for predicting reading behaviors, offering benefits to various stakeholders within the book ecosystem and fostering new applications for humanities research and cultural heritage valorization (Solimine / Zanchini 2020; Vivarelli 2018). Authors and publishers could use such tools to better understand and anticipate reading behaviors, improving text quality, usability, and effectiveness. Libraries, archives, and other cultural institutions could enhance the accessibility and valorization of their collections (Faggiolani et al. 2023), developing personalized and interactive reading paths, while readers could benefit from more engaging and tailored reading experiences, positively impacting the dissemination and understanding of cultural content. Furthermore, the above-mentioned correlation between triggers explicitly marked by readers and fixation points detected by the eye-tracker, if extended to LLM trigger predictions, might reinforce the interest of recent works exploring possible correlations between gaze patterns and transformer attention weights (Lopez-Cadorna et al. 2025), although such relationships remain interpretatively complex.

A comprehensive discussion of the results of the AI experiment will require a more detailed analysis than the one possible here, but the first results are very promising: we have observed a significant correlation between the exit points identified by the different LLMs used, as well as a significant correlation between the exit points identified by LLMs and those identified by human readers. Furthermore, at least some among the models used seem to be able – if properly prompted – to mimic the augmented reading behaviors of specific groups of readers.

References
  1. Allen, Graham (2021): Intertextuality. London: Routledge. 
  2. Baron, Naomi S. (2015): Words on Screen. The Fate of Reading in a Digital World. Oxford-New York: Oxford University Press. 
  3. Baron, Naomi S. (2021): How we read now. Oxford-New York: Oxford University press. 
  4. Ciotti, Fabio (2024): “Gli LLM come lettori modello artificiali” in: Di Silvestro, Antonio / Spampinato, Daria (eds.): Me.Te. Digitali. Mediterraneo in rete tra testi e contestiProceedings del XIII Convegno Annuale AIUCD2024, Catania, ottobre 2024: 342–344 <https://dx.doi.org/10.6092/unibo/amsacta/7927>
  5. Eco, Umberto (1979): Lector in fabula. Milano: Bompiani.  
  6. Faggiolani, Chiara / Federici, Alessandra / Quaglieri, Camilla (2023): “Biblioteche, infrastrutture culturali e polifunzionalità: una mappatura data driven”, in: AIB Studi, 63, 2: 245–262 <https://doi.org/10.2426/aibstudi-13883>   
  7. Faggiolani, Chiara / Verna, Lorenzo / Vivarelli, Maurizio (2017): “Text mining e network science per analizzare la complessità della lettura”, in: JLIS.it, 8:115–136 <https://doi.org/10.4403/jlis.it-12414
  8. Faggiolani, Chiara / Vivarelli, Maurizio (2016): Le reti della lettura. Tracce, modelli, pratiche del social reading. Milano: Editrice Bibliografica. 
  9. Franzmann, Bodo et al. (eds.) (1999): Handbuch Lesen. München: Saur.
  10. Iser, Wolfgang (1978): L'atto della lettura. Una teoria della risposta estetica. Italian trans., Bologna: Il Mulino. (Original work Der Akt des LesensThe Act of Reading, 1978). 
  11. Johns, Adrian (2023): The Science of Reading. Chicago: University of Chicago Press. 
  12. Landow, George P. (2006): Hypertext 3.0. Baltimore: Johns Hopkins University Press. 
  13. Lopez-Cadorna, Angela / Idesis, Sebastian / Barreda-Ángeles, Miguel /, Abadal, Sergi / Arapakis, Ioannis (2025): “OASST-ETC Dataset: Alignment Signals from Eye-tracking Analysis of LLM Responses”, in: Proceedings of the ACM on Human-Computer Interaction, 9, 3 <https://doi.org/10.1145/3725840>
  14. McGann, Jerome (2001): Radiant Textuality: Literature After the World Wide Web. New York: Palgrave.  
  15. Mason, Jessica (2019): Intertextuality in practice. Philadelphia: John Benjamins Publishing Company. 
  16. Mastrobattista, Ludovica (2024): “La lettura nell’era digitale: impatto sulla percezione del lettore”, in: AIB Studi64, 1: 27–39 <https://doi.org/10.2426/aibstudi-13994>
  17. Meyer, Timothy / Baker, Carolyn / Keefe, Jonathan (2025): “Enhancing thematic analysis with Local Large Language Models: A scientific evaluation of prompt engineering techniques”, In: Usability and User Experience, 194: 73–83 <https://doi.org/10.54941/ahfe1006669>
  18. Orr, Mary (2003): Intertextuality. Debates and Contexts. Cambridge: Polity. 
  19. Ranfa, Elena (2023): Sguardi sulla lettura. Percorsi tra le dimensioni del leggere. Manziana (Roma): Vecchiarelli Editore. 
  20. Rettberg, Scott (2018): Electronic Literature. Cambridge: Polity. 
  21. Roncaglia, Gino (2020): L’età della frammentazione: cultura del libro e scuola digitale. Bari-Roma: Laterza.
  22. Roncaglia, Gino (2022): “Letture aumentate, fra rete e intermedialità”, in: AIB Studi, 61, 3: 603–609 <https://doi.org/10.2426/aibstudi-13360
  23. Roncaglia, Gino (2023): L’architetto e l’oracolo. Forme di organizzazione digitale della conoscenza, da Wikipedia a ChatGPT. Bari-Roma: Laterza.
  24. Roncaglia, Gino (2025): Filosofia dell’IA. In Lezioni di filosofia. Milano: Corriere della Sera.  
  25. Salter, Anastasia / Moulthrop, Stuart (2021): Twining: Critical and Creative Approaches to Hypertext Narratives. Amherst: Amherst College Press.
  26. Solimine, Giovanni / Zanchini, Giorgio (2020): La cultura orizzontale. Bari-Roma: Laterza.
  27. Suleiman, Susan R. / Crosman, Inge (1980): The Reader in the Text. Essays on Audience and Interpretation. Oxford: Princeton University Press.
  28. Tabbi, Joseph (ed.) (2018): The Bloomsbury handbook of electronic literature. London: Bloomsbury.
  29. Vivarelli, Maurizio (2018): La lettura: Storie, teorie, luoghi. Milano: Editrice Bibliografica.  
  30. Wolf, Maryanne (2007): Proust and the Squid. The Story and Science of the Reading Brain. New York: HarperCollins.
  31. Wolf, Maryanne (2018). Reader, Come Home. New York: HarperCollins.