DH 2026

Daejeon, July 27–31

Poster

The Architecture of Fictional Conversation: Tracing Linguistic Alignment in Austen and Forster

Ioana-Roxana Boriceanu
University of Bucharest, Romania · ioana-roxana.boriceanu@s.unibuc.ro
Alina Iacob
University of Bucharest, Romania · ioana.iacob@s.unibuc.ro
Liviu P. Dinu
University of Bucharest, Romania; Human Language Technologies Research Center · ldinu@fmi.unibuc.ro

Abstract

This study examines whether linguistic coordination occurs in fictional dialogue. Applying LIWC-based convergence measures and network analysis to novels by Austen and Forster, we observe consistent alignment patterns across characters. The results indicate that coordination emerges even in imagined exchanges, pointing to a cognitive basis for how authors construct dialogue. We also note differences between authors, with Austen's networks showing higher reciprocity and cohesion than Forster's more fragmented patterns.

Introduction and Related Work

Linguistic coordination has been extensively studied in linguistics, cognitive science, and social psychology as a means of understanding how speakers adapt to one another during conversation. The phenomenon appears under various labels, including alignment (Pickering / Garrod 2004), accommodation (Giles et al. 1991), entrainment (Levitan / Hirschberg 2011), and linguistic style matching (Gonzales et al. 2010; Ireland / Henderson 2014), yet all describe the tendency for interlocutors to make their language more similar over the course of an exchange. Communication Accommodation Theory (Giles et al. 1991) emphasises social motivations, while the Interactive Alignment Model (Pickering / Garrod 2004) foregrounds automatic cognitive processes that support shared representations. Computational work has shown that coordination relates to engagement, cohesion, and power dynamics (Babcock et al. 2014; Danescu-Niculescu-Mizil et al. 2012), and can be both strategic and unconscious (Doyle / Frank 2016).

An important contribution to this area is the observation that linguistic coordination appears even in imagined conversations extracted from movie scripts (Danescu-Niculescu-Mizil / Lee 2011), suggesting it may stem from cognitive routines shaping dialogue representation rather than being driven exclusively by social factors.

This study analyses fictional dialogue in novels by Jane Austen and E. M. Forster to determine whether linguistic coordination occurs in literary texts, how patterns vary between authors, and how adaptation structures character interactions. We acknowledge that our corpus represents canonical British fiction, and coordination patterns may differ in other literary traditions. Forster explicitly described Austen as an influence (Colmer 1982), providing useful literary context for the comparison.

Data and Methods

Data

We use the Project Dialogism Novel Corpus (PDNC) (Vishnubhotla et al. 2022), a manually annotated collection of English novels providing speaker and addressee information for each quotation, along with character metadata such as gender and narrative importance. Quotations were cleaned by removing punctuation and special symbols and lowercasing the text.

Since most authors in the PDNC are represented by a single work, we focus on the novels by Jane Austen and E. M. Forster, the two authors with the largest coverage. Table 1 presents an overview of the corpus. From merged dialogue turns we created adjacent speaker pairs (A → B, B → A) and corresponding non-adjacent and randomised controls to distinguish local adaptation from broader stylistic similarity.

AuthorNovelCharsQuotesTokens
AustenEmma181,59376,141
Mansfield Park361,15758,121
Northanger Abbey2084228,227
Persuasion3548826,252
Pride and Prejudice741,27048,132
Sense and Sensibility231,08551,184
ForsterA Passage to India422,12337,088
A Room with a View631,63429,694
Howards End512,60646,709
Where Angels Fear to Tread181,00018,560
Austen total2066,435288,057
Forster total1747,363132,051

Table 1: Corpus overview: characters, quotations, and tokens per novel.

Methods

Our approach follows the convergence framework of Danescu-Niculescu-Mizil / Lee 2011, which measures how often a speaker's use of a linguistic feature is followed by the same feature in an interlocutor's reply. We analyse nine LIWC function-word families (Pennebaker et al. 2015; Ireland et al. 2011), typically processed non-consciously (Chung / Pennebaker 2011): articles, auxiliary verbs, conjunctions, adverbs, impersonal pronouns, negations, personal pronouns, prepositions, and quantifiers (566 lexemes total).

For each speaker pair A and B and each family f, convergence is computed as:

Figure 1: Convergence measure for each LIWC family.

Positive values indicate that B is more likely to use a feature after A has used it, signalling convergence. Negative values indicate divergence, while values near zero reflect no systematic adaptation.

To contextualise pairwise patterns, we construct convergence networks where nodes represent characters and directed, weighted edges encode mean convergence across the nine families. We analyse these networks using measures such as density, reciprocity, and assortativity to characterise how linguistic adaptation contributes to the structure of fictional interactions.

Experiments and Results

We applied our convergence and network analyses to the ten novels. Figure 2 presents convergence patterns for both authors. Austen's dialogues display clear convergence across most LIWC families, with adjacent turns producing higher adaptation than non-adjacent or random controls. Forster's novels also exhibit convergence, though in a more selective and uneven pattern.

(a)

(b)

(c)

(d)

Figure 2: Convergence patterns. Left column (a, c): conditional vs. baseline probability by LIWC family; asterisks indicate statistical significance (p < 0.05). Right column (b, d): convergence scores across conditions — adjacent pairs, non-adjacent pairs, and randomised controls. Top row (a, b): Austen; bottom row (c, d): Forster.

Gender effects follow earlier findings (Danescu-Niculescu-Mizil / Lee 2011): female characters tend to adapt more than male characters, while initiator gender shows no stable trend (Figure 3).

(a)

(b)

(c)

(d)

Figure 3: Gender effects. Left column (a, c): convergence by initiator gender. Right column (b, d): convergence by responder gender. Top row (a, b): Austen; bottom row (c, d): Forster.

Network analysis results appear in Table 2 and Figure 4. Austen's dialogue networks exhibit high reciprocity (0.88) and broad cohesion, with convergence distributed across many speakers. Forster's networks are more fragmented and context-dependent, with lower reciprocity (0.84). These patterns reveal meaningful differences in how each author constructs interactional structure in fiction.

MetricAustenForster
Nodes9574
Edges312237
Density0.0350.044
Reciprocity0.880.84
Positive edges (%)4854
Negative edges (%)3126
Gender assortativity−0.160.03

Table 2: Network metrics for the two corpora.

Figure 4: Sample convergence network from the Austen corpus (Pride and Prejudice).

Conclusions

This study demonstrates that linguistic convergence is not confined to natural conversation but also shapes fictional dialogue. Applying established coordination measures to novels by Austen and Forster, we find that alignment emerges even in imagined exchanges, suggesting that authors draw on cognitive routines underlying interaction. Differences between authors indicate that convergence interacts with narrative design and stylistic choices, offering a quantitative perspective on how social dynamics are encoded in fiction. A more detailed analysis, including per-novel breakdowns and community detection, is presented in Boriceanu et al. 2026.

Future work could extend this analysis to additional authors, genres, historical periods, and non-Western literary traditions, as well as incorporate variables such as class or social position. These findings position linguistic coordination as a useful tool for analysing style and character relations in literary texts.

Acknowledgements

This research was supported by the Ministry of Education and Research, CNCS-UEFISCDI, project SIROLA, number PN-IV-P1-PCE-2023-1701, within PNCDI IV, and by the project “Romanian Hub for Artificial Intelligence – HRIA”, Smart Growth, Digitization, and Financial Instruments Program, 2021–2027, MySMIS no. 334906.

References
  1. Babcock, Meghan J. / Ta, Vivian P. / Ickes, William (2014): “Latent semantic similarity and language style matching in initial dyadic interactions”, in: Journal of Language and Social Psychology 33.1: 78–88.
  2. Boriceanu, Ioana-Roxana / Iacob, Alina / Dinu, Liviu P. (2026): “Voices and Echoes in Fictional Dialogue: A Study of Linguistic Coordination in Literary Texts”, in: Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). Palma, Mallorca: ELRA: 2581–2593.
  3. Chung, Cindy / Pennebaker, James (2011): “The psychological functions of function words”, in: Social Communication. Psychology Press: 343–359.
  4. Colmer, John (1982): “Marriage and Personal Relations in Forster’s Fiction”, in: EM Forster: Centenary Revaluations. Springer: 113–123.
  5. Danescu-Niculescu-Mizil, Cristian / Lee, Lillian (2011): “Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs”, in: Proceedings of the 2nd Workshop on Cognitive Modeling and Computational Linguistics: 76–87.
  6. Danescu-Niculescu-Mizil, Cristian / Lee, Lillian / Pang, Bo / Kleinberg, Jon (2012): “Echoes of power: Language effects and power differences in social interaction”, in: Proceedings of the 21st International Conference on World Wide Web: 699–708.
  7. Doyle, Gabriel / Frank, Michael C. (2016): “Investigating the sources of linguistic alignment in conversation”, in: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers): 526–536.
  8. Giles, Howard / Coupland, Nikolas / Coupland, Justine (1991): “Accommodation theory: Communication, context, and consequence”, in: Contexts of Accommodation: Developments in Applied Sociolinguistics 1: 1–68.
  9. Gonzales, Amy L. / Hancock, Jeffrey T. / Pennebaker, James W. (2010): “Language style matching as a predictor of social dynamics in small groups”, in: Communication Research 37.1: 3–19.
  10. Ireland, Molly E. / Slatcher, Richard B. / Eastwick, Paul W. / Scissors, Lauren E. / Finkel, Eli J. / Pennebaker, James W. (2011): “Language style matching predicts relationship initiation and stability”, in: Psychological Science 22.1: 39–44.
  11. Ireland, Molly E. / Henderson, Marlone D. (2014): “Language style matching, engagement, and impasse in negotiations”, in: Negotiation and Conflict Management Research 7.1: 1–16.
  12. Levitan, Rivka / Hirschberg, Julia Bell (2011): “Measuring acoustic-prosodic entrainment with respect to multiple levels and dimensions”.
  13. Pennebaker, James W. / Boyd, Ryan L. / Jordan, Kayla / Blackburn, Kate (2015): The Development and Psychometric Properties of LIWC2015.
  14. Pickering, Martin J. / Garrod, Simon (2004): “Toward a mechanistic psychology of dialogue”, in: Behavioral and Brain Sciences 27.2: 169–190.
  15. Vishnubhotla, Krishnapriya / Hammond, Adam / Hirst, Graeme (2022): “The Project Dialogism Novel Corpus: A dataset for quotation attribution in literary texts”, in: Proceedings of the Thirteenth Language Resources and Evaluation Conference: 5838–5848.