DH 2026

Daejeon, July 27–31

Wed, July 2909:00–10:30S108Grand Ballroom
Short Paper

PERson Counting: Operationalizing a Theory of Character in an NLP Pipeline

Mark Andrew Algee-Hewitt
Stanford University, United States of America · malgeehe@stanford.edu
Unjoo Oh
Stanford University, United States of America · ujoh@stanford.edu
Jessica Monaco
Stanford University, United States of America · jmmonaco@stanford.edu
Gabi Keane
Stanford University, United States of America · gab03@stanford.edu
Emil Miao Wang
Stanford University, United States of America · wshimiao@stanford.edu

Computational studies of literature that touch on issues of characters within fictional texts have both opened new avenues to the study of character at the scale of corpora and simultaneously have revealed gaps between the computational and critical identification of characters. By treating characters as a measurable quantity within texts, our overall project leverages the advantages of this approach to explore the relationship between the size of character systems and the length of the texts that contain them. In doing so, however, we seek to redress a crucial difference between how literary characters are typically counted in computational analyses and how they are theorized critically.

We take it as a given that longer texts can contain more unique characters, as there is simply more space to introduce readers to more people. But, unlike many other facets of literary texts, the relationship between text length and number of unique characters is not linear. Previous work (Algee-Hewitt, Mukamal, and Porter, 2021) has demonstrated that although the average novel contains more unique characters than a short story, the rate at which characters are introduced falls sharply between short stories (15 characters per 10,000 words) and novels (6 characters per 10,000 words). Figure 1 demonstrates graphically that as the length of a novel increases the number of character introductions follow a power law distribution: increasing rapidly at shorter lengths but slowing down appreciably at larger text sizes (R-Squared=0.57).

Figure 1: Novel length vs number of unique characters in a corpus of 9449 English language 20th century novels.

This project extends this previous effort in two ways. First, our overall project seeks to explore this relationship across a much larger corpus of texts, including long-form prose fiction in English, French, German, and Spanish, as well as texts across a larger historical continuum. Our goal is to discover the upper bounds for the number of characters that can be introduced in a novel and whether this changes between time, form, and language.

Second, current state-of-the-art analyses of characters in fiction largely depends on entity recognition within an NLP pipeline, most commonly Bamman et al.’s BookNLP (2019). Although such a born-literary pipeline radically improves both entity detection and, more importantly, co-reference resolution, studies that use this method typically equate entities labeled as “PER” (or persons) and characters. In other words, they begin with the assumption that all persons mentioned in a text are characters within that fictional world (see, for example, Underwood et al. 2018, Soni et al., 2023, Piper 2024). While this can be a useful heuristic, it fails to account for the reader’s experience of character and introduces noise in the form of a wide variety of questionable entities into the analysis. For example, our first test of this pipeline compared five expert reader’s free annotation of character in the Virginia Woolf short story “Kew Gardens” with unique entities identified using BookNLP’s co-reference resolution (figure 2).

Figure 2: Number of character annotations of five expert readers vs unique entities found via BookNLP in “Kew Gardens.”

Without guiding rules for annotating character, the expert readers demonstrated low inter-annotator reliability, identifying between 9 and 34 characters. BookNLP, however, counted 81 entities, including “the spirits of the dead,” “mermaids,” “women drowned at sea,” and the sententious use of “one.” None of these entities (all of which were mentioned purely symbolically in the text) were tagged by human readers as “characters.” This is a significant problem for our proposed project as the prevalence of anthropomorphized symbols and abstract metaphors is confounded with style and period such that entity counts vary widely according to these variables rather than the actual presence of characters in the text.

Researchers have sought to solve this challenge by limiting the analysis to “main” characters: for example, Lucy et al. (2025) track sentences containing entities whose co-reference ID is among the top 3% of all character mentions (after hand-correcting BookNLP’s co-reference resolution). Other researchers choose to adopt different approaches, particularly for non-English corpora, including fine-tuning both a BERT-based classifier and a LLM to identify character properties (Pagel and Reiter 2025). The presence of characters within a given text segment, however, was again based on a model-based NER algorithm. These studies artificially limit the number of characters they quantify by either ignoring minor characters or focusing only on characters with identifiable properties: in neither case does the solution answer the needs of this study, which aims to count all characters, whether main or minor.

The challenge of identifying characters computationally has been recognized by prominent scholars working between literary theory and computational literary study. For example, Fotis Jannidis has identified the incompatibility between theories of characters that insist all characters are language-based, and those that resolve character to the representation of persons in the minds of readers (Eder et al. 2010). While the first suggests the tractability of the problem to computational methods, the second does not. Even for traditional literary criticism, the definition of character is elided in favor of studies of characterization or character systems (Lynch 1998, Vermeule 1998, Culler, 1975). Even Alex Woloch, in The Once vs the Many classifies characters but refrains from identify their ontological properties as characters (2003).

Our short presentation of this in-progress project will demonstrate our solution to this challenge. Our pipeline also begins with an NER pass based on BookNLP which we, following Lucy et al. also correct for non-collapsed, co-reference entities. Through a process of annotation, however, we have developed a set of guidelines for identifying characters, both human and non-human, that we implement within BookNLP (figure 3).

Figure 3: Annotation rules for identifying characters in a fictional text

Beginning with a corrected set of PER co-references from BookNLP, we implement rules 2 and 3 using POS tags and dependency annotations, rule 4 based on collapsed counts, and rules 1, 5, and 6 using both a model and vocabulary-based approach to the type’s attachments in the dependency parse. Following our case study above, we implemented these rules and automatically identified 9 characters in “Kew Gardens”, equivalent to the mode of our reader annotations. Although more complex, we argue that this pipeline yields results that are not only better in keeping with reader annotations, but it also operationalizes literary theoretical approaches to character within a computational pipeline.

References
  1. Algee-Hewitt, Mark / Mukamal, Anna / Porter, JD (2023): “The Affordances of Mere Length: Computational Approaches to Short Story Analysis.” in The Cambridge Companion to the American Short Story ed. MJ Collins and G Jones. Cambridge University Press: 341-357.
  2. Bamman, David (2024). “Born Literary Natural Language Processing.” Debates in the Digital Humanities: Computational Humanities. Ed. Lauren Tilton, David Mimno, and Jessica Marie Johnson. University of Minnesota Press: np.
  3. Culler, Jonathan (1975): Structuralist Poetics: Structuralism, Linguistics, and the Study of Literature. Routledge.
  4. Eder, Jens / Jannidis, Fotis / Schneider, Ralf (2010): “Introduction.” Characters in Fictional Worlds: Understanding Imaginatory Beings in Literature, Film, and Other Media ed. J Eder, F. Jannidis, R. Schneider. De Gruyter. 3-67.
  5. Lucy, L. / Griffiths, C. / Ying, C. / Kim-Ebio, J.J. / Baur, S./ Levine, S. / Eberhardt, J.L. / Bamman, D. / Demszky, D. (2025): “Racial and Ethnic Representation in Literature Taught in US High Schools.” Journal of Cultural Analytics, 10, 1.
  6. Lynch, D. (1998): The economy of character: novels, market culture, and the business of inner meaning. University of Chicago Press.
  7. Pagel, J. / Reiter, N. (2025): “Automatic Detection and Classification of Literary Character Properties in German Narratives.” Anthology of Computers and the Humanities, 3:1494-1509.
  8. Piper, A. (2024): “What do characters do? The embodied agency of fictional characters.” Journal of Computational Literary Studies2(1).
  9. Soni, S. / Sihra, A. / Wilkins, M. / Evans, E. / Bamman, D. (2023): “Grounding Characters and Places in Narrative Texts.” Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Volume 1: Long Papers. 11723-11736.
  10. Underwood, T. / Bamman, D. / Lee, S. (2018): “The transformation of gender in English-language Fiction.” Cultural Analytics. Feb. 13, 2018. DOI: 10.22148/16.019.
  11. Vermeule, B. (2010): Why do we care about literary characters?. Johns Hopkins University Press.
  12. Woloch, Alex (2003): The One vs the Many: Minor Characters and the Space of the Protagonist. Princeton University Press.