DH 2026

Daejeon, July 27–31

Poster

The GOLEM Annotated Corpus for Computational Literary Studies

Federico Pianzola
University of Groningen, Netherlands, The · f.pianzola@rug.nl

We present a multilingual dataset of fiction in 7 languages, enriched with manual annotation about characters and events in the stories. The dataset includes documents of diverse lengths, ranging from 500 to 17k words, with a total of around 100k words per language (Table 1). We collected texts in seven languages (Chinese, Dutch, English, Indonesian, Italian, Korean, and Spanish). The corpus was compiled from three online fanfiction platforms: Archive of Our Own (AO3; all languages), Postype (Korean), and Wattpad (Dutch and Indonesian). To maintain a balanced and diverse corpus across languages, we sampled works along multiple dimensions, including fandom, worldbuilding settings, narrative perspective, dialogue/narration ratio, genre, and popularity (kudos/hits ratio).

Manual annotation has been done by at least two trained annotators (native speakers) per language, with an additional curation step performed by an expert team of native speakers. The annotations cover the following aspects:

  • characters’ mentions and coreference resolution (van Cranenburgh et al. 2026);
  • characters’ traits (biography, physical, personality), emotions, and mental states (Yang and Pianzola 2025);
  • dialogue markers, including speaker and addressee (Papay and Padó 2020);
  • relationships between characters (free text);
  • event segmentation (Gius and Vauth 2022, Visser Solissa et al. 2025);
  • event types (Gius and Vauth 2022, Visser Solissa et al. 2025).

This ensemble of annotations can be used to train and evaluate language models for a variety of tasks in computational literary studies. It is the first resource of this kind that also extensively covers non-European languages.

Table 1. Summary statistics of the annotated corpus

ChineseDutchEnglishIndonesianItalianKoreanSpanishTotal
Tokens104,772119,989134,348132,719120,05890,031125,086827,003
Sentences3,9079,9538,83610,9016,1489,6866,22155,652
Mentions8,18816,42718,51119,36814,4919,97515,722102,682
Entities5709137416845526678164,943
Tokens/sentence26.812.115.212.219.59.320.114.9
Mentions/tokens0.0780.1370.1380.1460.1210.1110.1260.124
Mentions/entity14.3617.9924.9828.3226.2514.9619.2720.77

Existing related datasets that focus only on one or few languages are:

  • Chinese: CSI (Yu et al. 2022), SPC (Chen et al. 2023), SIG (Su et al. 2024);
  • Dutch: dutchcoref (van Cranenburgh 2019), OpenBoek (van Cranenburgh and van Noord 2022), (van Cranenburgh and van den Berg 2023);
  • English: QuoteLi3 (Muzny et al. 2017), LitBank and BookNLP (Bamman et al. 2020), RiQuA (Papay and Padó 2020), PDNC (Vishnubhotla et al. 2022), Cuesta-Lazaro et al. 2022, AustenAlike (Yang and Anderson 2024), Piper et al. 2024, DramaCV (Michel et al. 2024), TF-CRL (Dobson et al. 2025), BOOKCOREF (Martinelli et al. 2025);
  • French: Durandard et al. 2023, BookNLP-fr (Mélanie-Becquet et al. 2024);
  • German: DROC (Krug et al. 2017), Redewiedergabe (Brunner et al. 2021), LLpro (Ehrmanntraut et al. 2023), GerFuN (Hilger and Ehrmanntraut 2026);
  • Japanese: Gobara et al. 2025.
  • Korean: Koconovel (Kim et al. 2024), Jang and Jung 2024;
  • Portuguese: Taggus (Canário et al. 2026);
  • Multilingual: ELTeC (Byszuk et al. 2020).

By extending the range of languages covered by resources for computational literary studies, we hope that more researchers will engage in comparative analysis of literature on a large scale.

All stories included in our dataset are either works' for which authors voluntarily waived copyright or we obtained explicit permission from authors to analyze and share the data after anonymization. The repository https://github.com/GOLEM-lab/GOLEM-multilingual-annotated-corpus contains thefull annotated corpus available in various formats, the annotation guidelines used for each task, as well as reports on the challenges identified during the annotation and curation.

References
  1. Bamman, David / Lewke, Olivia / Mansoor, Anya (2020): "An annotated dataset of coreference in English literature", in: Proceedings of the Twelfth Language Resources and Evaluation Conference. Marseille: European Language Resources Association: 44–54. <https://aclanthology.org/2020.lrec-1.6/>.
  2. Brunner, Annelen / Tu, Ngoc Duyen Tanja / Weimer, Lukas / Jannidis, Fotis (2020): "To BERT or not to BERT: Comparing contextual embeddings in a deep learning architecture for the automatic recognition of four types of speech, thought and writing representation", in: Proceedings of the 5th Swiss Text Analytics Conference (SwissText) & 16th Conference on Natural Language Processing (KONVENS). CEUR Workshop Proceedings 2624. Zurich. <https://ceur-ws.org/Vol-2624/>.
  3. Byszuk, Joanna / Woźniak, Michał / Kestemont, Mike / Leśniak, Albert / Łukasik, Wojciech / Šeļa, Artjoms / Eder, Maciej (2020): "Detecting direct speech in multilingual collection of 19th century novels", in: Proceedings of the LREC 2020 Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2020). Marseille: 100–104. <https://aclanthology.org/2020.lt4hala-1.16/>.
  4. Canário, Tiago G. / Duarte, Carlos / Pinheiro, Flávio L. / Pereira, José L. (2026): "Taggus: An Automated Pipeline for the Extraction of Characters' Social Networks from Portuguese Fiction Literature", in: Social Network Analysis and Mining 16, 40. DOI: 10.1007/s13278-026-01584-6.
  5. Chen, Yue / He, Tianwei / Zhou, Hongbin / Gu, Jia-Chen / Lu, Heng / Ling, Zhen-Hua (2023): "Symbolization, Prompt, and Classification: A Framework for Implicit Speaker Identification in Novels", in: Findings of the Association for Computational Linguistics: EMNLP 2023. Singapore: 3455–3467. DOI: 10.18653/v1/2023.findings-emnlp.225.
  6. Cranenburgh, Andreas van (2019): "A Dutch coreference resolution system with an evaluation on literary fiction", in: Computational Linguistics in the Netherlands Journal 9: 27–54. <https://clinjournal.org/clinj/article/view/91>.
  7. Cranenburgh, Andreas van / Berg, Frank van den (2023): "Direct Speech Quote Attribution for Dutch Literature", in: Proceedings of the 7th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature. Dubrovnik: Association for Computational Linguistics: 45–62. DOI: 10.18653/v1/2023.latechclfl-1.6.
  8. Cranenburgh, Andreas van / Noord, Gertjan van (2022): "OpenBoek: A Corpus of Literary Coreference and Entities with an Exploration of Historical Spelling Normalization", in: Computational Linguistics in the Netherlands Journal 12: 235–251. <https://clinjournal.org/clinj/article/view/157>.
  9. Cranenburgh, Andreas van / Yang, Xiaoyan / Alvanita / Di Domenico, Cecilia Nicole / Ferragud, Maria / Graciotti, Arianna / Kim, Byungjun / Park, Seonyeong / Visser Solissa, Noa / Zhou, Xiaoyu / Pianzola, Federico (2026): "GOLEMcoref: A Multilingual Coreference Dataset of Fiction", in: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics. San Diego: Association for Computational Linguistics.
  10. Cuesta-Lazaro, Carolina / Prasad, Animesh / Wood, Trevor (2022): "What does the sea say to the shore? A BERT based DST style approach for speaker to dialogue attribution in novels", in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Dublin: Association for Computational Linguistics: 5820–5829. DOI: 10.18653/v1/2022.acl-long.400.
  11. Ehrmanntraut, Anton / Konle, Leonard / Jannidis, Fotis (2023): "LLpro: A Literary Language Processing Pipeline for German Narrative Texts", in: Proceedings of the 19th Conference on Natural Language Processing (KONVENS 2023). Ingolstadt: Association for Computational Linguistics: 28–39. <https://aclanthology.org/2023.konvens-main.3/>.
  12. Gius, Evelyn / Vauth, Michael (2022): "Towards an Event Based Plot Model. A Computational Narratology Approach", in: Journal of Computational Literary Studies 1, 1. DOI: 10.48694/jcls.110.
  13. Gobara, Seiji / Kamigaito, Hidetaka / Watanabe, Taro (2025): "Speaker Identification and Dataset Construction Using LLMs: A Case Study on Japanese Narratives", in: Proceedings of the 7th Workshop on Narrative Understanding. Albuquerque, NM: Association for Computational Linguistics: 97–119. DOI: 10.18653/v1/2025.wnu-1.17.
  14. Hilger, Agnes / Ehrmanntraut, Anton (2026): "Coreference Resolution for Full German Novels using Large Language Models", in: CCLS2026 Conference Preprints 5, 1. DOI: 10.26083/tuda-7983.
  15. Jang, Woori / Jung, Seohyon (2024): "Evaluating LLM Performance in Character Analysis: A Study of Artificial Beings in Recent Korean Science Fiction", in: Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities. Miami: Association for Computational Linguistics: 339–351. DOI: 10.18653/v1/2024.nlp4dh-1.34.
  16. Kim, Kyuhee / Lee, Surin / Lee, Sangah (2024): "KoCoNovel: Annotated dataset of character coreference in Korean novels". ArXiv preprint. DOI: 10.48550/arXiv.2404.01140.
  17. Koolen, Corina / Cranenburgh, Andreas van (2018): "Blue eyes and porcelain cheeks: Computational extraction of physical descriptions from Dutch chick lit and literary novels", in: Digital Scholarship in the Humanities 33, 1: 59–71. DOI: 10.1093/llc/fqx016.
  18. Krug, Markus / Weimer, Lukas / Reger, Isabella / Macharowsky, Luisa / Feldhaus, Stephan / Puppe, Frank / Jannidis, Fotis (2017): "Description of a corpus of character references in German novels – DROC [Deutsches ROman Corpus]". DARIAH-DE Working Papers 27. Göttingen: DARIAH-DE. URN: urn:nbn:de:gbv:7-dariah-2018-2-9.
  19. Martinelli, Giuliano / Bonomo, Tommaso / Huguet Cabot, Pere-Lluís / Navigli, Roberto (2025): "BOOKCOREF: Coreference Resolution at Book Scale", in: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna: Association for Computational Linguistics: 24526–24544. DOI: 10.18653/v1/2025.acl-long.1197.
  20. Mélanie-Becquet, Frédérique / Barré, Jean / Seminck, Olga / Plancq, Clément / Naguib, Marco / Pastor, Martial / Poibeau, Thierry (2024): "BookNLP-fr, the French Versant of BookNLP. A Tailored Pipeline for 19th and 20th Century French Literature", in: Journal of Computational Literary Studies 3, 1: 1–34. DOI: 10.48694/jcls.3924.
  21. Michel, Gaspard / Epure, Elena V. / Hennequin, Romain / Cerisara, Christophe (2024): "Improving Quotation Attribution with Fictional Character Embeddings", in: Findings of the Association for Computational Linguistics: EMNLP 2024. Miami, FL: Association for Computational Linguistics: 12723–12735. DOI: 10.18653/v1/2024.findings-emnlp.744.
  22. Muzny, Grace / Fang, Michael / Chang, Angel / Jurafsky, Dan (2017): "A Two-stage Sieve Approach for Quote Attribution", in: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. Valencia: Association for Computational Linguistics: 460–470. <https://aclanthology.org/E17-1044/>.
  23. Papay, Sean / Padó, Sebastian (2020): "RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text", in: Proceedings of the Twelfth Language Resources and Evaluation Conference. Marseille: European Language Resources Association: 835–841. <https://aclanthology.org/2020.lrec-1.104/>.
  24. Piper, Andrew / Xu, Michael / Ruths, Derek (2024): "The Social Lives of Literary Characters: Combining citizen science and language models to understand narrative social networks", in: Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities. Miami: Association for Computational Linguistics: 472–482. DOI: 10.18653/v1/2024.nlp4dh-1.45.
  25. Su, Zhenlin / Xu, Liyan / Xu, Jin / Li, Jiangnan / Huangfu, Mingdu (2024): "SIG: Speaker identification in literature via prompt-based generation", in: Proceedings of the AAAI Conference on Artificial Intelligence 38, 17: 19035–19043. DOI: 10.1609/aaai.v38i17.29870.
  26. Vishnubhotla, Krishnapriya / Hammond, Adam / Hirst, Graeme (2022): "The Project Dialogism Novel Corpus: A Dataset for Quotation Attribution in Literary Texts", in: Proceedings of the Thirteenth Language Resources and Evaluation Conference. Marseille: European Language Resources Association: 5838–5848. <https://aclanthology.org/2022.lrec-1.628/>.
  27. Visser Solissa, Noa / Cranenburgh, Andreas van / Pianzola, Federico (2025): "Event Detection between Literary Studies and NLP. A Survey, a Narratological Reflection, and a Case Study", in: Journal of Computational Literary Studies 4, 1. DOI: 10.48694/jcls.4215.
  28. Yang, Funing / Anderson, Carolyn Jane (2024): "Evaluating Computational Representations of Character: An Austen Character Similarity Benchmark", in: Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities. Miami: Association for Computational Linguistics: 17–30. DOI: 10.18653/v1/2024.nlp4dh-1.3.
  29. Yang, Xiaoyan / Pianzola, Federico (2025): "Fans Reconstruct Heroes: Modeling Fictional Characters in Participatory Culture", in: Semantic Web – Interoperability, Usability, Applicability. <https://www.semantic-web-journal.net/system/files/swj3885.pdf>.
  30. Yu, Dian / Zhou, Ben / Yu, Dong (2022): "End-to-End Chinese Speaker Identification", in: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Seattle: Association for Computational Linguistics: 2274–2285. DOI: 10.18653/v1/2022.naacl-main.165.