Daejeon, July 27–31
This study examines whether linguistic coordination occurs in fictional dialogue. Applying LIWC-based convergence measures and network analysis to novels by Austen and Forster, we observe consistent alignment patterns across characters. The results indicate that coordination emerges even in imagined exchanges, pointing to a cognitive basis for how authors construct dialogue. We also note differences between authors, with Austen's networks showing higher reciprocity and cohesion than Forster's more fragmented patterns.
Linguistic coordination has been extensively studied in linguistics, cognitive science, and social psychology as a means of understanding how speakers adapt to one another during conversation. The phenomenon appears under various labels, including alignment (Pickering / Garrod 2004), accommodation (Giles et al. 1991), entrainment (Levitan / Hirschberg 2011), and linguistic style matching (Gonzales et al. 2010; Ireland / Henderson 2014), yet all describe the tendency for interlocutors to make their language more similar over the course of an exchange. Communication Accommodation Theory (Giles et al. 1991) emphasises social motivations, while the Interactive Alignment Model (Pickering / Garrod 2004) foregrounds automatic cognitive processes that support shared representations. Computational work has shown that coordination relates to engagement, cohesion, and power dynamics (Babcock et al. 2014; Danescu-Niculescu-Mizil et al. 2012), and can be both strategic and unconscious (Doyle / Frank 2016).
An important contribution to this area is the observation that linguistic coordination appears even in imagined conversations extracted from movie scripts (Danescu-Niculescu-Mizil / Lee 2011), suggesting it may stem from cognitive routines shaping dialogue representation rather than being driven exclusively by social factors.
This study analyses fictional dialogue in novels by Jane Austen and E. M. Forster to determine whether linguistic coordination occurs in literary texts, how patterns vary between authors, and how adaptation structures character interactions. We acknowledge that our corpus represents canonical British fiction, and coordination patterns may differ in other literary traditions. Forster explicitly described Austen as an influence (Colmer 1982), providing useful literary context for the comparison.
We use the Project Dialogism Novel Corpus (PDNC) (Vishnubhotla et al. 2022), a manually annotated collection of English novels providing speaker and addressee information for each quotation, along with character metadata such as gender and narrative importance. Quotations were cleaned by removing punctuation and special symbols and lowercasing the text.
Since most authors in the PDNC are represented by a single work, we focus on the novels by Jane Austen and E. M. Forster, the two authors with the largest coverage. Table 1 presents an overview of the corpus. From merged dialogue turns we created adjacent speaker pairs (A → B, B → A) and corresponding non-adjacent and randomised controls to distinguish local adaptation from broader stylistic similarity.
| Author | Novel | Chars | Quotes | Tokens |
| Austen | Emma | 18 | 1,593 | 76,141 |
| Mansfield Park | 36 | 1,157 | 58,121 | |
| Northanger Abbey | 20 | 842 | 28,227 | |
| Persuasion | 35 | 488 | 26,252 | |
| Pride and Prejudice | 74 | 1,270 | 48,132 | |
| Sense and Sensibility | 23 | 1,085 | 51,184 | |
| Forster | A Passage to India | 42 | 2,123 | 37,088 |
| A Room with a View | 63 | 1,634 | 29,694 | |
| Howards End | 51 | 2,606 | 46,709 | |
| Where Angels Fear to Tread | 18 | 1,000 | 18,560 | |
| Austen total | 206 | 6,435 | 288,057 | |
| Forster total | 174 | 7,363 | 132,051 |
Table 1: Corpus overview: characters, quotations, and tokens per novel.
Our approach follows the convergence framework of Danescu-Niculescu-Mizil / Lee 2011, which measures how often a speaker's use of a linguistic feature is followed by the same feature in an interlocutor's reply. We analyse nine LIWC function-word families (Pennebaker et al. 2015; Ireland et al. 2011), typically processed non-consciously (Chung / Pennebaker 2011): articles, auxiliary verbs, conjunctions, adverbs, impersonal pronouns, negations, personal pronouns, prepositions, and quantifiers (566 lexemes total).
For each speaker pair A and B and each family f, convergence is computed as:
Figure 1: Convergence measure for each LIWC family.
Positive values indicate that B is more likely to use a feature after A has used it, signalling convergence. Negative values indicate divergence, while values near zero reflect no systematic adaptation.
To contextualise pairwise patterns, we construct convergence networks where nodes represent characters and directed, weighted edges encode mean convergence across the nine families. We analyse these networks using measures such as density, reciprocity, and assortativity to characterise how linguistic adaptation contributes to the structure of fictional interactions.
We applied our convergence and network analyses to the ten novels. Figure 2 presents convergence patterns for both authors. Austen's dialogues display clear convergence across most LIWC families, with adjacent turns producing higher adaptation than non-adjacent or random controls. Forster's novels also exhibit convergence, though in a more selective and uneven pattern.
(a) | (b) |
(c) | (d) |
Figure 2: Convergence patterns. Left column (a, c): conditional vs. baseline probability by LIWC family; asterisks indicate statistical significance (p < 0.05). Right column (b, d): convergence scores across conditions — adjacent pairs, non-adjacent pairs, and randomised controls. Top row (a, b): Austen; bottom row (c, d): Forster.
Gender effects follow earlier findings (Danescu-Niculescu-Mizil / Lee 2011): female characters tend to adapt more than male characters, while initiator gender shows no stable trend (Figure 3).
(a) | (b) |
(c) | (d) |
Figure 3: Gender effects. Left column (a, c): convergence by initiator gender. Right column (b, d): convergence by responder gender. Top row (a, b): Austen; bottom row (c, d): Forster.
Network analysis results appear in Table 2 and Figure 4. Austen's dialogue networks exhibit high reciprocity (0.88) and broad cohesion, with convergence distributed across many speakers. Forster's networks are more fragmented and context-dependent, with lower reciprocity (0.84). These patterns reveal meaningful differences in how each author constructs interactional structure in fiction.
| Metric | Austen | Forster |
| Nodes | 95 | 74 |
| Edges | 312 | 237 |
| Density | 0.035 | 0.044 |
| Reciprocity | 0.88 | 0.84 |
| Positive edges (%) | 48 | 54 |
| Negative edges (%) | 31 | 26 |
| Gender assortativity | −0.16 | 0.03 |
Table 2: Network metrics for the two corpora.
Figure 4: Sample convergence network from the Austen corpus (Pride and Prejudice).
This study demonstrates that linguistic convergence is not confined to natural conversation but also shapes fictional dialogue. Applying established coordination measures to novels by Austen and Forster, we find that alignment emerges even in imagined exchanges, suggesting that authors draw on cognitive routines underlying interaction. Differences between authors indicate that convergence interacts with narrative design and stylistic choices, offering a quantitative perspective on how social dynamics are encoded in fiction. A more detailed analysis, including per-novel breakdowns and community detection, is presented in Boriceanu et al. 2026.
Future work could extend this analysis to additional authors, genres, historical periods, and non-Western literary traditions, as well as incorporate variables such as class or social position. These findings position linguistic coordination as a useful tool for analysing style and character relations in literary texts.
This research was supported by the Ministry of Education and Research, CNCS-UEFISCDI, project SIROLA, number PN-IV-P1-PCE-2023-1701, within PNCDI IV, and by the project “Romanian Hub for Artificial Intelligence – HRIA”, Smart Growth, Digitization, and Financial Instruments Program, 2021–2027, MySMIS no. 334906.