Daejeon, July 27–31
Characters in narratives are characterized in various ways, including their inner and outer features, such as attire, height, and personality as well as patterns in the language attributed to them. Additionally, the language they use plays an important role in characterizing them (Page 1988, Kinsui 2003). This paper examines the words of characters in The Tale of Genji, one of the oldest extant Japanese novels written during the Heian period (794-1192), with the aim of demonstrating how characters’ words contribute to their characterization. Quantitative approaches have recently been introduced into the study of The Tale of Genji as well as other classical Japanese literature of this period, and the number of such studies remains limited (e.g., Tsuchiayama & Murakami 2013, Xu et al. 2025). Therefore, most previous studies addressing character creation in the tale rely on qualitative analysis. For instance, they investigate the characterization of each character by analyzing their roles in the story (Nakajima 2004). However, recent quantitative studies, such as Kondo (2005) and Takeuchi and Ogiso (2026), indicate that characters’ words also play a pivotal role in shaping characters by identifying gender distinctions embodied in linguistic choices made by characters. Kondo, for example, conducts an n-gram analysis of the short poems composed by the characters and discusses how their diction is carefully chosen to represent gender ideology of the Heian period. This finding challenges the traditional assumption, as Endo (1997) states, that classical Japanese lacked gender distinctions in grammar or vocabulary, suggesting instead that characterization may be achieved through subtle but statistically significant patterns in their diction. Yet her analysis leaves other forms of characters’ words unexamined. Thus, no quantitative investigation has yet demonstrated how all forms of characters’ words contribute to their characterization in The Tale of Genji. This study, therefore, addresses this gap by examining all forms of characters’ words by using quantitative methods across the narrative.
In this study, we conduct a multivariate analysis of emotive adjectives produced by characters across all forms of expressions, such as conversation, inner monologue, and written matters, focusing on female characters in The Tale of Genji. Building on previous findings in Takeuchi and Ogiso (2026), which identifies dominant patterns, or gender ideologies, in the use of emotive adjectives among characters classified by gender and social group, this study demonstrates that linguistic choices contribute systematically to characterization in The Tale of Genji based on Heian-period gender ideology. It also provides new empirical evidence for discussions of gendered language and characterization in classical Japanese literature.
Data
We utilize the data for The Tale of Genji included in the Heian Period Series of the Corpus of Historical Japanese (The National Institute for Japanese Language and Linguistics 2016). The tale contains 445,711 tokens, 151,199 of which are categorized as characters’ words. We recently expanded it by adding various speaker information, including gender, name, and utterance categorization (Takeuchi et al. 2022). This study investigates the 34 high-frequency emotive adjectives examined in Takeuchi and Ogiso (2026), as shown in Table 1.
Table 1: The targeted emotive adjectives
Method and Procedure
This study involves two different multivariate analyses: a hierarchical agglomerative analysis and a correlation analysis. First, we perform a hierarchical agglomerative analysis, which detects similar data points and group them into clusters by finding patterns and connections in datasets. This enables us to discern female characters who use emotive adjectives similarly. Then, we perform a correlation analysis, which determines if there is a correlation between two variables by providing a quantifiable measure of the strength and the nature of the association between two or more variables. This analysis allows us to determine how similarly each group uses these adjectives compared to the dominant patterns of the character categories found in Takeuchi and Ogiso (2026): the nun (Nun) and the female layperson (F-LP). These two categories used for correlation analysis are derived from the aggregate linguistic patterns of all characters belonging to these social groups across the entire story, which provide a baseline of group-level gender and social norms. We utilize Python to process the data and perform the analyses.
Analysis and Results
In The Tale of Genji, female characters tend to produce few words. Therefore, we investigate 12 female characters who produce the emotive adjectives more than 20 times in total and their use of these adjectives. Table 2 presents the description of the emotive adjectives produced by the female characters.
Table 2
The data above is analyzed using the hierarchical agglomerative cluster technique. Parameters for cluster identification are as follows: Pearson’s correlation, ratio, Ward’s method. Figure 1 shows the resulting tree plot. The y-axis represents dissimilarity or the distance between clusters/variables (1- correlation coefficient). We cut the tree at a distance of 1.0, where the correlation becomes zero.
Figure 1: Tree plot
Four large clusters are generated. Cluster A includes three low-class characters, two of whom are nuns. Cluster B includes three noble characters. Two of them renounce the world later in the story while the other is raised by a hermit-like father. Cluster C includes three noble characters, two of whom marry the protagonist, Genji. Cluster D includes three noble characters who are related to the imperial family by marriage or blood.
We then conduct a correlation analysis to measure how strongly each cluster is associated with the language use of the aggregate language patterns of the two female character categories: the nun characters (Nun) and the female layperson characters (F-LP). During the Heian period, renouncing the world impacted people's lives in various ways, including style of dress, relationships, and roles in life. Thus, this social group distinction played a pivotal role, reflected in their language use (Takeuchi & Ogiso 2026). Figure 2 presents the resulting matrix.
Figure 2: Correlation matrix
Cluster A has a strong positive correlation with Nun (r = .7839), which is a stronger association than F-LP (r = .5432), while Cluster C shows almost no correlation with Nun (r = -.0302) yet a medium correlation with F-LP (r = .4940). Cluster D has a medium correlation with F-LP (r = .5632) and a relatively weak correlation with Nun (r = .3980). Cluster B shows medium correlations with Nun (r = .5374) and F-LP (r = .6688).
As seen above, Cluster B is moderately correlated with the two categories, and the difference between the correlation coefficients is the smallest. In other words, this cluster more or less shows a similar tendency to both categories in the use of the emotive adjectives. To identify the emotive adjectives associated with the dominant patterns, we generate correlation scatter plots.
Figure 3: Correlation scatter plots
Most emotive adjectives are plotted relatively close to the regression lines. However, the adjectives kurushii, ui, and tsurai are notable outliers, plotted farther away. This statistical deviation shows that the characters in Cluster B tend to express pain through their words more frequently than the aggregate baseline for their social group. This aligns with their situations in the story. Indeed, they all experience suffering in their relationships with male characters. For example, when Ukifune is distraught over her complicated relationships and her circumstances, she uses these specific emotive markers to express her suffering, such as “uki mono” (a thing of sorrow). Of these adjectives, ui and tsurai are gendered words, reflecting the gender ideology of the Heian period (Takeuchi & Ogiso 2026). Tsurai is associated with the male layperson characters, while ui is associated with the female layperson characters. Thus, the noble women in Cluster B violate the norm by using male-appropriate adjectives to a significant degree. This deviation illustrates how specific linguistic choices contribute to shaping these characters' identities within the narrative.
Conclusion
In this paper, we present how characters’ words contribute to character construction by conducting multivariate analyses of high-frequency emotive adjectives. These analyses suggest that the prominent linguistic choices in wording not only reflect characters’ situations in the story but also indicate their qualities as characters. Specifically, identifying noble women’s use of male-appropriate adjectives demonstrates that characterization in The Tale of Genji often functions by deviating from established Heian-period gender ideologies. For our next steps, we will investigate the prominence of each cluster and each character in their words to deepen our understanding of how they are shaped by their words, which would otherwise go undetected. This will improve our understanding of the characters as well as the classical Japanese language.