DH 2026

Daejeon, July 27–31

Thu, July 3013:40–15:10S094108
Long Paper

“Letting the Cat out of the Bag”: Towards the Automatic Detection of Animal Metaphors in Literary Texts

Katrin Rohrbacher
FAU Erlangen-Nürnberg, Germany · katrin.rohrbacher@fau.de
Andreas Wagner
FAU Erlangen-Nürnberg, Germany · andreas.w.wagner@fau.de
Michaela Mahlberg
FAU Erlangen-Nürnberg, Germany · michaela.mahlberg@fau.de

Introduction

Animal metaphors are a central but conceptually elusive feature of literary language, often used to characterize humans, encode social dynamics, or signal moral and cultural values. An example is “Another step, you beast, and husband or no husband, I’ll kill you!” which situates social power relations in animal imagery. Despite their interpretive importance, animal metaphors are difficult to define and identify systematically across large text collections. This paper explores their detection using methods from Natural Language Processing (NLP).

For code and data of the project, see, https://github.com/gi47huwi/AniMe.

Metaphors are widely recognized as complex, context-sensitive, and theoretically contested phenomena, with extensive study in linguistics (Semino 2008; Pragglejaz 2007; Fauconnier and Turner 2003; Goatly 1997; Lakoff and Johnson 1980), philosophy (Sontag 1989; Ricœur 1993; Davidson 1978; Blumenberg 1960), and literary scholarship (Wellek 1949; Jakobson 1956). Computational research has developed methods to detect metaphors more broadly, focusing on figurative language at scale (Ptiček and Dobša 2023; Tsvetkov et al. 2013; Stefanowitsch and Gries 2006), but quantitative studies specifically on animal metaphors are rare. Examples in Cultural Analytics include analyses of animacy (Häußler et al. 2024; Karsdorp et al. 2015) or anthropomorphization (Hara and Koda 2020).

Even when metaphor is not the main focus of a study, its pervasive nature makes it relevant across different areas. For instance, in their study of biodiversity in creative literature, Langer et al. (2021) extract terms referring to animals using a dictionary matching approach. Piper (2022) replicated their study, using a supervised machine learning approach (Book NLP’s supersense tags), to identify occurrences of words referring to animals and other biological entities. These approaches provide insights into animal reference distribution but do not distinguish literal from figurative use. Studies specifically targeting animal metaphors remain limited. The closest precedent for combining manual annotation with quantitative metaphor analysis in literary texts is Caracciolo et al. (2019), who systematically code human–nonhuman metaphors, including animal imagery, across three novels. While Caracciolo et al. (2019) demonstrate a close-reading approach, we make a proposal for coding animal metaphors using a transformer-based classifier that will allow us to cover a corpus of over 4,000 works.

In our study, we focus on animal metaphors, human-related or not, and explore their automatic detection by training a classifier on a large, diverse set of texts to identify them at scale. We treat animal metaphor detection as a binary task and adopt a broad working definition: an ‘animal’ term is metaphorical if it creates a contrast between its literal meaning and its contextual meaning (Pragglejaz 2007). Additionally, we work at the sentence level, in contrast to more common token-level approaches to general metaphor detection (Tong et al. 2021). Similes are treated as a subset of metaphorical expressions due to their role in conveying meaning beyond the literal.

Our definition does not distinguish between novel and conventional or ‘dead’ metaphors. While this distinction is beyond the scope of the present study, it remains an important direction for future work.

Animal Metaphor Detection as Sentence-Level Annotation

The Corpus Data

Our corpus consists of 4,246 works of English prose fiction spanning the years 1780 to 1930 and covering twelve genres, retrieved from Project Gutenberg using metadata provided by the Gender Novel Corpus (Tagliaferri et al. 2019). It includes a broad range of narrative subgenres (speculative fiction, crime fiction, fairy tales, and realist novels among them), making it well-suited for studying how animal metaphor use varies across generic and historical contexts.

Our corpus matches the Gender Novel Corpus in terms of text selection; the metadata provided by Tagliaferri et al. (2019) was used to identify and retrieve the texts from Project Gutenberg.

We applied BookNLP (Bamman et al., 2014)

https://github.com/booknlp/booknlp

to the corpus, using its supersense tagger to identify nouns tagged as noun.animal in the WordNet supersense taxonomy as candidate animal terms. From these, we randomly sampled a candidate set of 5,500 sentences for annotation, preserving genre and temporal variation across the sample.

Manual Annotation

The first and second authors independently annotated a random subset of 500 sentences drawn from the candidate set to establish inter-annotator agreement. A sentence was coded as “metaphor” if it contained an animal metaphor according to our working definition, and as “no metaphor” if it did not. For each candidate sentence, a preceding and following sentence were shown to provide context (see Figure 1, where the term pet has been tagged as noun.animal).

Figure 1. Label Studio’s interface for annotation.

While many metaphorical mentions were straightforward to annotate, a subset of cases constituted recurring “edge cases”, where literal and figurative readings overlap. In practice, the contrast between literal and non-literal use typically entails that animal qualities are attributed to a human referent or abstract entity; edge cases tend to be those where this transfer is absent or unclear. These edge cases fall into four types.

The first type involves passages where literal animals are embedded in figurative frames, e.g., an elephant trampling a cabin “like an ant-hill”. Here, the noun elephant is used in its literal meaning and the simile conveys scale and does not attribute animal qualities to humans.

The second type involves similes that do not attribute animal qualities to a human referent but contribute to the characterisation of a wider context:

“Now, out here in this desolate place … the very frogs tired of their chanting … these two far, deep-toned syllables seemed like a human voice.”

The frogs are physically present in the scene and it is the sound, not a human referent, that is being characterized via the animal. Such cases are nonetheless prone to misclassification, as the simile structure and animal term together closely resemble the surface form of a true animal metaphor.

The third type involves metaphors that target collective or non-human entities rather than individual humans (e.g., “The middle class is a wobbly little lamb between a lion and a tiger”), where the target falls outside the core case of individual human referents, introducing uncertainty about whether the annotation criteria apply.

The fourth type involves syntactically marked constructions such as enumerations, which introduce a further kind of ambiguity. In “A wild animal, a man, a snake, might be in hiding,” the parallel listing of a human alongside animals invites the inference that the man shares their qualities, without making that mapping explicit. Unlike a simile or direct predication, the figurative meaning is implied through syntactic parallelism alone, making it difficult to classify under a definition that requires a recoverable contrast between literal and contextual meaning.

Inter-annotator agreement was high in terms of percentage agreement (≈89%), though Krippendorff’s α was 0.58 once chance agreement was taken into account. The comparatively lower α largely reflects the strong imbalance between metaphorical and non-metaphorical instances, as well as the interpretive nature of metaphor identification, particularly in borderline cases like those described above. Agreement measures of this range align with comparative studies on general metaphor detection in computational linguistics.

See, for instance, Shutova (2017), where agreement between two annotators on identifying linguistic metaphors reached κ = 0.62, or Maudslay et al. (2024), who report κ = 0.51 for metaphor identification.

Following consensus discussions to resolve disagreements, the annotators each independently annotated a further portion of the remaining candidate set. The process resulted in a final annotated dataset of 1,823 sentences, of which 361 were tagged as “metaphor” and 1,462 as “no metaphor” (see Figure 2).

Figure 2. Distribution of publication dates across the annotated dataset, by metaphor label.

Model and Training

Given the class imbalance between “no metaphor” and “metaphor” sentences, we randomly sampled the majority class during training to reduce overfitting. A RoBERTa-based binary classification model (“roberta-base”)

https://huggingface.co/FacebookAI/roberta-base

was fine-tuned with a learning rate of 1e-5 on 70% of the annotated sentences and evaluated on a held-out validation set of 30% over five epochs (see Table 1). Our best model achieved an F1 of 0.72 for metaphor detection, comparable to other NLP studies on computational metaphor processing (see, for instance, Mao et al. 2023; Maudslay et al. 2020; Gao et al. 2018). While these existing quantitative approaches focus on general metaphor detection, directly applying them to our task is not straightforward: they are primarily trained and evaluated on news and academic corpora, and may not generalise well to the distinct linguistic and stylistic properties of literary fiction. We therefore opted for a dedicated classifier trained on literary data.

Table 1. Performance measures of the RoBERTa classification model. Precision, Recall, and F1-score are reported per class (Non-metaphor, Metaphor) as well as macro and weighted averages. Support indicates the number of instances per class in the validation set.

Validation

To understand the model’s limitations and inform future improvements, we manually inspected misclassifications. Across the four edge case types discussed above, the classifier performed well, correctly handling the majority of complex and ambiguous cases. The one exception was the enumerative construction (“A wild animal, a man, a snake, might be in hiding”), where the classifier predicted “no metaphor” — though given the implicit nature of the mapping, this is better understood as reflecting the genuine difficulty of the case than as a clear misclassification.

The main source of false positives was simile-like structures without an accompanying animal metaphor, pointing to a limitation in distinguishing formal comparison from metaphorical function. An example from the corpus is “…with no more warning than the sound of a swamp-bird’s flight, was like a nightmare,” where the bird is used literally and the simile maps the thought onto a nightmare rather than attributing animal qualities to a human referent. Future training data should therefore include more non-metaphorical similes to enhance this differentiation, as well as more targeted examples of rare constructions.

Conclusion

This paper demonstrates that animal metaphors can be detected automatically at the sentence level using a relatively small but carefully annotated training set. Despite the small training set and the heterogeneous nature of metaphor occurrences, the classifier performs reasonably well and aligns with existing metaphor detection benchmarks.

Future work should expand training data to better cover rare edge cases and include more non-metaphorical similes. While the classifier has been evaluated on a held-out sample drawn from the full corpus, testing on a larger sample would further assess its robustness across the full genre and temporal range. Once sufficiently robust, the model can be applied to investigate how animal metaphor use varies across genres, time periods, and author demographics — questions that have so far been inaccessible at corpus scale due to the difficulty of distinguishing figurative from literal animal reference. For fields such as animal studies and ecocriticism, which rely heavily on close reading of individual texts, this opens up new possibilities. It becomes possible to trace how cultural attitudes toward the nonhuman world and social values encoded in animal imagery are distributed across the full breadth of a literary tradition.

References
  1. Bamman, David / Underwood, Ted / Smith, Noah A. (2014): "A Bayesian Mixed Effects Model of Literary Character." In: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics: 370–379. https://doi.org/10.3115/v1/P14-1035.
  2. Blumenberg, Hans (1960): Paradigmen zu einer Metaphorologie. 2. Auflage. Suhrkamp Studienbibliothek 10. Suhrkamp.
  3. Caracciolo, Marco / Ionescu, Andrei / Fransoo, Ruben (2019): “Metaphorical Patterns in Anthropocene Fiction.” In: Language and Literature 28, 3: 221–40. https://doi.org/10.1177/0963947019865450.
  4. Davidson, Donald (1978): “What Metaphors Mean.” In: Critical Inquiry 5, 1: 31–47.
  5. Fauconnier, Gilles / Turner, Mark (2003): “Conceptual Blending, Form and Meaning.” In: Recherches en Communication 19. https://doi.org/10.14428/rec.v19i19.48413.
  6. Gao, Ge / Choi, Eunsol / Choi, Yejin / Zettlemoyer, Luke (2018): “Neural Metaphor Detection in Context.” In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics: 607–13. https://doi.org/10.18653/v1/D18-1060.
  7. Goatly, Andrew (1997): The Language of Metaphors. London / New York: Routledge.
  8. Hara, Kanae / Koda, Naoko (2020): “Quantitative Analysis of Anthropomorphic Animals in Picture Books.” In: International Journal of Literature and Arts 8, 6: 308. https://doi.org/10.11648/j.ijla.20200806.11.
  9. Häußler, Julian / von Keitz, Janis / Gius, Evelyn (2024): Animacy in German Folktales. CHR 2024, Aarhus, Denmark. https://ceur-ws.org/Vol-3834/paper90.pdf. 1023–1036.
  10. Jakobson, Roman / Halle, Morris (2020 [1956]): Fundamentals of Language. 2nd rev. ed. Janua Linguarum. Series Minor 1. De Gruyter Mouton. https://doi.org/10.1515/9783110889611.
  11. Karsdorp, Folgert / van der Meulen, Marten / Meder, Theo / van den Bosch, Antal (2015): “Animacy Detection in Stories.” In: OASIcs, Volume 45, CMN 2015 45: 82–97. https://doi.org/10.4230/OASICS.CMN.2015.82.
  12. Lakoff, George / Johnson, Mark (1980): Metaphors We Live By. Chicago: University of Chicago Press.
  13. Langer, Lars / Burghardt, Manuel / Borgards, Roland / Böhning-Gaese, Katrin / Seppelt, Ralf / Wirth, Christian (2021): "The Rise and Fall of Biodiversity in Literature: A Comprehensive Quantification of Historical Changes in the Use of Vernacular Labels for Biological Taxa in Western Creative Literature." In: People and Nature 3, 5: 1093–1109. https://doi.org/10.1002/pan3.10256
  14. Mao, Rui / Li, Xiao / He, Kai / Ge, Mengshi / Cambria, Erik (2023): "MetaPro Online: A Computational Metaphor Processing Online System." In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). Toronto, Canada: Association for Computational Linguistics: 127–135. https://doi.org/10.18653/v1/2023.acl-demo.12.
  15. Maudslay, Rowan Hall / Pimentel, Tiago / Cotterell, Ryan / Teufel, Simone (2020): “Metaphor Detection Using Context and Concreteness.” In: Proceedings of the Second Workshop on Figurative Language Processing: 221–26. https://doi.org/10.18653/v1/2020.figlang-1.30.
  16. Maudslay, Rowan Hall / Teufel, Simone / Bond, Francis / Pustejovsky, James (2024): “ChainNet: Structured Metaphor and Metonymy in WordNet.” arXiv:2403.20308. https://doi.org/10.48550/arXiv.2403.20308.
  17. Piper, Andrew (2022): “Biodiversity Is Not Declining in Fiction.” In: Journal of Cultural Analytics 7, 3. https://doi.org/10.22148/001c.38739.
  18. Pragglejaz Group (2007): “MIP: A Method for Identifying Metaphorically Used Words in Discourse.” In: Metaphor and Symbol 22, 1: 1–39. https://doi.org/10.1080/10926480709336752.
  19. Ptiček, Martina / Dobša, Jasminka (2023): “Methods of Annotating and Identifying Metaphors in the Field of Natural Language Processing.” In: Future Internet 15, 6: 201. https://doi.org/10.3390/fi15060201.
  20. Ricœur, Paul (1993): The Rule of Metaphor. University of Toronto Romance Series 37. Toronto: University of Toronto Press.
  21. Semino, Elena (2008): Metaphor in Discourse. Cambridge: Cambridge University Press.
  22. Shutova, Ekaterina (2017): “Annotation of Linguistic and Conceptual Metaphor.” In: Ide, Nancy / Pustejovsky, James (eds.): Handbook of Linguistic Annotation. Dordrecht: Springer: 1073–1100.
  23. Sontag, Susan (1989): Illness as Metaphor; and, AIDS and Its Metaphors. New York: Picador USA.
  24. Stefanowitsch, Anatol / Gries, Stefan Thomas (2006): Corpus-Based Approaches to Metaphor and Metonymy. Trends in Linguistics 171. Berlin: Mouton de Gruyter.
  25. Tagliaferri, Lisa et al. (2019): Computational Reading of Gender in Novels. https://github.com/dhmit/gender_novels_site [02.12.2025].
  26. Tkachenko, Maxim / Malyuk, Mikhail / Holmanyuk, Andrey / Liubimov, Nikolai (2020–2025): Label Studio: Data Labeling Software. https://github.com/HumanSignal/label-studio.
  27. Tong, Xiaoyu / Shutova, Ekaterina / Lewis, Martha (2021): “Recent Advances in Neural Metaphor Processing.” In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: 4673–86. https://doi.org/10.18653/v1/2021.naacl-main.372.
  28. Tsvetkov, Yulia / Mukomel, Elena / Gershman, Anatole (2013): “Cross-Lingual Metaphor Detection Using Common Semantic Features.” In: Proceedings of the First Workshop on Metaphor in NLP. Atlanta, Georgia, 13 June 2013: 45–51.
  29. Wellek, René / Warren, Austin (1963 [1949]): Theory of Literature. Third edition. Peregrine Books. London: Penguin Books.