Daejeon, July 27–31
This paper presents a reader-centered framework for evaluating automatic layout segmentation in webtoons to support computational analysis or “distant viewing” (Arnold / Tilton 2019). Webtoons are a form of online comics designed for an infinite vertical canvas, typically consumed through continuous scrolling on mobile devices (Jin 2019). This mode of presentation departs fundamentally from the page-based logic that underpins most established theories of comics layout (Berube et al. 2024; Fatha / Mansoor 2021). Rather than being organized into static, multi-panel pages, webtoons unfold as a continuous vertical narrative stream. While the screen provides a fixed viewing frame, temporal and spatial organization emerges through the reader’s scrolling behavior, which shapes both navigation and narrative pacing (Berube et al. 2024; Fatha / Mansoor 2021; Jin 2019).
Existing frameworks for analyzing comics layout—largely developed for printed media—primarily focus on spatial organization, visual grouping, and grid-based panel structures (Bateman et al. 2016; Chavanne 2015; Cohn 2013; Groensteen 2013; Peeters 2021; Mikkonen / Lautenbacher 2019). Their explanatory power is therefore limited when applied to non-page-based storytelling. Similarly, current automatic layout analysis models, such as YOLO-based object detection models and the Segment Anything Model (SAM), tend to equate visually separable regions with structurally meaningful narrative units, an assumption embedded in both their training objectives and evaluation metrics (Vivoli et al. 2024). This assumption, however, does not fully correspond to how narrative segmentation emerges during digital reading. Research on digital comics reading further suggests that scrolling gestures are not merely navigational actions but also part of how readers structure narrative absorption, making them an important empirical basis for segmentation (Berube et al. 2024).
In the absence of stable page or panel hierarchies, a central challenge lies in defining what constitutes a meaningful narrative unit in webtoon reading and how detecting such units can be operationalized for automatic segmentation. This challenge highlights a deeper issue: current automated tools rely on assumptions that are structurally incompatible with the narrative mechanics of the infinite canvas. Meaningful evaluation of segmentation models consequently requires a reader-centered account of narrative unit formation—one that integrates micro-scrolling behavior with semiotic analysis to establish empirically and theoretically valid criteria against which computational outputs can be meaningfully compared.
Building on this premise, this study addresses a central question: how should the effectiveness of automatic layout segmentation methods be evaluated when narrative units in webtoons are not only defined by visual boundaries, but also emerge from reader-centered, semiotically structured narrative flow? To tackle this question, we focus on three interrelated issues: 1. What kinds of units can be considered narratively meaningful segments during the reading of webtoons? 2. To what extent do existing automatic segmentation models align with the narrative units as perceived by readers? 3. How might annotation standards grounded in reader behavior and semiotic analysis reshape both the evaluation criteria and training practices of automatic segmentation methods? To address these questions, we adopt a three-step research design. We first recruit five participants to read webtoons of different narrative types on mobile devices, recording their reading behavior. These behavioral observations, combined with semiotic analysis, inform our definition of the layout units to be segmented. Using this reader-centered framework, we then manually annotate a webtoon corpus. Finally, we apply and compare several widely used automatic segmentation models against this annotated reference. A more detailed explanation is provided below.
To identify narratively meaningful breakpoints in webtoon reading, we first conducted a small-scale exploratory reading study with five participants engaging with webtoons of different genres and visual styles on mobile devices. Screen recordings allowed us to capture fine-grained interaction traces, including brief pauses, slight reversals, accelerated scrolling, and zooming in specific scenes. We interpret these micro-scrolling gestures as rhythmic markers that often coincide with plot progression, emotional shifts, or changes in visual focus, and therefore as indicators of narrative flow. While exploratory and qualitative in scope, the study does not aim at statistical generalization. Rather, it provides an interpretive basis for examining how segmentation emerges during actual reading. Informed by these behavioral observations and refined through semiotic analysis of visual and narrative cues, we define layout units for annotation in a way that remains closely aligned with reading experience. Annotation guidelines were iteratively refined through researcher discussion to establish interpretive consistency within a reader-centered segmentation framework. Building on this foundation, we created a diverse webtoon corpus for methodological exploration and comparative analysis. The corpus comprises several webtoon episodes spanning a range of genres and visual styles, ensuring variation in pacing, transition structure, and visual organization. Using the segmentation criteria established above, three researchers manually annotated the layout structures within the corpus and resolved disagreements through iterative discussion. This step allows us to assess the operational feasibility and interpretive adequacy of reader-oriented narrative units across different webtoon formats while also providing a reference point against which automated segmentation outputs can be evaluated.
Following corpus annotation, we apply several widely used automatic segmentation approaches—including YOLO-based object detection models and SAM—to the same set of webtoon episodes in order to generate machine-defined layout segmentations. These models primarily rely on learned visual representations that encode cues such as edges, color regions, and contrast to identify candidate structural units. Their outputs can therefore be understood as reflecting the default interpretations of webtoon layout produced by current vision-driven segmentation paradigms. To evaluate the applicability of these automatic segmentation methods to webtoon narrative flow, we directly compare model outputs with the narrative units identified through reader behavior and semiotically informed annotation. Rather than relying on conventional visual metrics such as intersection over union (IoU) or bounding-box overlap (Cheng et al. 2021 ; Everingham et al. 2010), our comparison focuses on whether the two forms of segmentation point to similar narrative transitions, rhythmic shifts, or meaning-bearing boundaries. In other words, we examine whether computational segmentation approximates the narrative structure recognized by readers during actual reading, rather than simply assessing its precision in detecting visual borders.
Preliminary comparisons reveal several recurrent patterns of bias in automatic segmentation models when applied to webtoons. First, models tend to over-segment at points of explicit visual discontinuity while overlooking implicit transitions shaped by narrative pacing or emotional buildup. Second, sequences of continuous images that function as coherent narrative units for readers are often fragmented into multiple isolated segments. Third, narrative signals associated with textual emphasis, color orchestration, or scrolling rhythm are only partially captured by visually driven segmentation approaches. The recurrence of these patterns across different webtoon genres and visual styles suggests that they are not isolated errors, but stem from a fundamentally visual understanding of narrative units embedded in current models. Rather than proposing an alternative automatic segmentation model, this study argues that segmentation in webtoons is inherently interpretive and cannot be evaluated independently of reader experience and narrative logic (Bateman / Wildfeuer 2014; D’Ignazio / Klein 2020). By demonstrating how small-scale, reader-centered data production and annotation can support a more context-sensitive evaluation framework, this research contributes a methodological perspective for assessing automatic segmentation beyond visual accuracy alone. More broadly, it invites reflection on the use of AI-based tools in digital narrative research and offers methodological insights for their responsible evaluation and application within digital humanities. The study therefore proposes a reader-centered framework for evaluating computational segmentation in webtoons that understands segmentation not simply as a visual task, but as a narrative and interpretive phenomenon.