Daejeon, July 27–31
We present a multilingual dataset of fiction in 7 languages, enriched with manual annotation about characters and events in the stories. The dataset includes documents of diverse lengths, ranging from 500 to 17k words, with a total of around 100k words per language (Table 1). We collected texts in seven languages (Chinese, Dutch, English, Indonesian, Italian, Korean, and Spanish). The corpus was compiled from three online fanfiction platforms: Archive of Our Own (AO3; all languages), Postype (Korean), and Wattpad (Dutch and Indonesian). To maintain a balanced and diverse corpus across languages, we sampled works along multiple dimensions, including fandom, worldbuilding settings, narrative perspective, dialogue/narration ratio, genre, and popularity (kudos/hits ratio).
Manual annotation has been done by at least two trained annotators (native speakers) per language, with an additional curation step performed by an expert team of native speakers. The annotations cover the following aspects:
This ensemble of annotations can be used to train and evaluate language models for a variety of tasks in computational literary studies. It is the first resource of this kind that also extensively covers non-European languages.
Table 1. Summary statistics of the annotated corpus
| Chinese | Dutch | English | Indonesian | Italian | Korean | Spanish | Total | |
| Tokens | 104,772 | 119,989 | 134,348 | 132,719 | 120,058 | 90,031 | 125,086 | 827,003 |
| Sentences | 3,907 | 9,953 | 8,836 | 10,901 | 6,148 | 9,686 | 6,221 | 55,652 |
| Mentions | 8,188 | 16,427 | 18,511 | 19,368 | 14,491 | 9,975 | 15,722 | 102,682 |
| Entities | 570 | 913 | 741 | 684 | 552 | 667 | 816 | 4,943 |
| Tokens/sentence | 26.8 | 12.1 | 15.2 | 12.2 | 19.5 | 9.3 | 20.1 | 14.9 |
| Mentions/tokens | 0.078 | 0.137 | 0.138 | 0.146 | 0.121 | 0.111 | 0.126 | 0.124 |
| Mentions/entity | 14.36 | 17.99 | 24.98 | 28.32 | 26.25 | 14.96 | 19.27 | 20.77 |
Existing related datasets that focus only on one or few languages are:
By extending the range of languages covered by resources for computational literary studies, we hope that more researchers will engage in comparative analysis of literature on a large scale.
All stories included in our dataset are either works' for which authors voluntarily waived copyright or we obtained explicit permission from authors to analyze and share the data after anonymization. The repository https://github.com/GOLEM-lab/GOLEM-multilingual-annotated-corpus contains thefull annotated corpus available in various formats, the annotation guidelines used for each task, as well as reports on the challenges identified during the annotation and curation.