Daejeon, July 27–31
Theoretical Frameworks
This study is anchored in three complementary cognitive and pedagogical frameworks, each aligned with a distinct research variable to build a comprehensive model for AI-mediated learning.
1. Aptitude–Treatment Interaction (ATI). This framework offers a foundation for scrutinizing the variable of instructional modality (reading versus listening) and its interplay with modality preference. The ATI model maintains that instructional effectiveness is fundamentally linked to the fit between the instruction type and the characteristics of the individual learner. Prior investigation has demonstrated that learners exhibiting a visual preference often show stronger comprehension in reading scenarios over auditory ones, simultaneously reporting a lower cognitive load. Consequently, the ATI model guides our analysis of how individual preferences mediate affective and cognitive engagement across auditory and visual instructional formats.
2. Levels of Processing (LoP) Theory. The LoP theory serves as the theoretical core for assessing the effects of both content elaboration (summaries versus full text) and playback speed. LoP proposes that the persistence of memory traces is directly correlated with the depth of initial cognitive effort, meaning that deep semantic engagement leads to superior recall compared to shallow encoding. This principle is employed to evaluate whether the simplified nature of AI-generated summaries results in less effective processing than full texts, or whether simplification facilitates deeper engagement by reducing initial cognitive demands. Furthermore, LoP directs our study of playback speed, assessing how the temporal aspect of audio consumption influences processing depth and efficiency.
3. Zone of Proximal Development (ZPD) Vygotsky's ZPD provides a developmental perspective concerning the utility of content elaboration. ZPD defines the space between a learner's independent ability and their potential achievement with appropriate support or scaffolding. We suggest that complex, full-length texts may exceed the current ZPD for learners with limited prior exposure. Therefore, AI-generated summaries may serve as effective pedagogical supports, modifying the learning material to bring it within the learner’s accessible zone, thus enabling deeper processing and facilitating skill progression. The combined application of ZPD, ATI, and LoP establishes a comprehensive perspective for this research.
Methodology and Work in Progress
This doctoral work employs a rigorous mixed-methods approach comprising two separate quantitative experiments followed by a qualitative phase. We are currently recruiting approximately 210 adult participants through Prolific, ensuring they are native English speakers without known sensory impairments.
Experiment 1: Modality and Content Elaboration. This experiment focuses on the interaction between the two independent variables: instructional modality (reading vs. listening) and content elaboration (full text vs. AI-generated summary). The experimental structure involves a 2x2 design, wherein participants are randomly assigned to one of four delivery conditions: reading the full text, listening to the full text, reading a summary, or listening to a summary. Each condition is planned to include 30 participants, totaling 120 for this phase.
Prior to the learning phase, participants complete a short survey to determine their modality preference (ATI framework). Following the timed learning process (to measure objective efficiency), participants complete an objective knowledge construction test and a self-reported questionnaire assessing perceived learning, enjoyment, and cognitive load. Analyses of variance will be the primary statistical method used.
Experiment 2: Playback Speed. This phase is designed to specifically investigate how the temporal component affects audio learning. The content is standardized (an AI-generated summary), while listening speed is varied across three distinct conditions: standard speed (1.0×), moderately increased speed (1.5×), and high speed (2.0×). This phase includes 90 participants (30 in each condition). Outcomes measured are consistent with the first experiment, including objective knowledge, affective responses, and efficiency.
Qualitative Complementary Phase. The final phase involves semi-structured interviews with a sample of 35 participants (5 from each of the seven experimental groups). The purpose is to enrich the quantitative findings with subjective insights, exploring participants’ self-reported comprehension, strategies for engagement, and their ease or difficulty experienced during processing.
Contribution and Anticipated Results
This research contributes theoretically by significantly advancing the application of ATI, LoP, and ZPD models to AI-supported digital discourse and learning systems. By rigorously examining how content complexity and temporal factors interact with learner characteristics, the study provides a nuanced understanding of cognitive outcomes and how AI impacts digital discourse in learning context.
On a practical level, the forthcoming findings will provide empirically supported principles intended to guide the development of adaptive learning software and inform the Digital Humanities community on the appropriate use of summarized versus full content, and the optimal adjustment of audio speed based on learner processing capacities.
Keywords: AI in Education, Audio Learning, Levels of Processing (LoP), Digital Pedagogy, Aptitude-Treatment Interaction (ATI), Learning Efficiency, Summarization