DH 2026

Daejeon, July 27–31

Poster

Visualising Dante’s Semantics of Fire through Lexicon Structure and Embedding-Based Similarity across the Divine Comedy

Ruoci Song
Cambridge University, United Kingdom · rs2268@cam.ac.uk

Fire in Dante’s Divine Comedy is a motif with a visible theological arc and a difficult computational form. It begins as punitive combustion in Inferno, becomes purgative ordeal in Purgatorio, and is progressively absorbed into light, ardour, and intellectual desire in Paradiso. Existing Dante criticism has described this transformation with great subtlety, especially in work on the semantic field of fire and heat (Pertile 1991). This poster foregrounds a complementary question: how can a literary motif whose meaning moves between literal flame and metaphorical ardour be modelled as a graded textual phenomenon beyond binary keyword search? A line such as “Al mio ardor fuor seme le faville” (Purg. XXI, 94) is a useful motivating case: it contains recognisable fire vocabulary, yet its force depends on poetic memory, desire, and local context.

This poster presents a lexicon-driven and embedding-based model for visualising the fire motif across the 14,233 lines of the Petrocchi edition of the Divine Comedy. The corpus is “enriched” only in a structural and documentary sense: each line is associated with cantica, canto, line number, and local textual context. The fire lexicon is manually constructed from corpus-attested forms, checked against philological discussion, and organised into three operational tiers, without relying on a general semantic network. CORE_FIRE includes literal combustion terms such as foco/fuoco (“fire”), fiamma (“flame”), incendio (“fire/conflagration”), favilla (“spark”), and physical uses of ardere (“to burn”); ABSTRACT_FIRE includes ardore (“ardour/heat”), scintilla (“spark”), sfavilla (“sparkles”), and figurative forms of accendere (“to kindle”); FALSE_POSITIVE covers proper names or forms such as Fiamminghi or Focaccia that contain fire-like strings without belonging to the motif. This answers a central interpretive problem: the lexicon is authorial and corpus-constrained, yet its tiering remains an explicit scholarly decision.

The model combines four components in a Flame Intensity Score (FIS): a tiered lexical score L, an embedding similarity score S, a local contextual score C, and a disambiguation penalty D for pseudo-fire. Sentence embeddings are produced with multilingual-e5-large (Wang et al. 2024), using normalised line vectors. S measures cosine similarity to a prototype vector built from strongly literal fire passages. C uses a three-line window, consisting of the previous, current, and following line, chosen as a compact approximation of local poetic context and aligned with verse-level reading practice. The weights are heuristic and deliberately transparent: they were set after pilot inspection of known fire loci and borderline cases, then kept fixed so that parameter sensitivity remains visible. This also makes the score readable in a poster setting, where viewers can see how lexical evidence, semantic similarity, context, and penalty interact. The score is therefore presented as a heuristic lens for interpretation, with no claim to ground-truth classification.

The poster’s main contribution lies in coordinating several visual views of the same score, so that the modelling process remains inspectable. Canto-level curves and a cantica–canto heatmap show expected concentrations around Inferno XIV-XVII and XXVI, the purgative threshold of Purgatorio XXVII, and a more diffuse paradisal pattern in which fire merges with light and ardour. The model also highlights less obvious boundary cases, where isolated fire lexemes fail to create a sustained cluster, or where abstract ardour locally reactivates the fire field. A UMAP projection of the scored verses offers a complementary semantic map: Inferno forms a dense punitive cluster, Purgatorio functions as a transitional bridge, and Paradiso disperses into a constellation of luminous and affective fire. The poster therefore combines a compact pipeline diagram with heatmaps, curves, and UMAP views, so that technical choices and interpretive consequences can be read together. These views are intended to send viewers back to the poem, making visible where quantitative density, lexical tiering, and local context agree or diverge.

The poster makes the source of the lexicon, the meaning of structural metadata, the size of the contextual window, and the status of the weights explicit. It also situates the work within distributional semantics and computational literary studies on topics and motifs (Navarro-Colorado 2018; van Zundert et al. 2022), while preserving the philological specificity of the Dante case. A lightweight implementation context is available through the research page at https://anticafiamma.it/research/fiamma, where the visualisations can be inspected alongside the textual evidence. The broader aim is to offer a reusable, inspectable workflow for motif analysis in historical literary texts, with fire in the Divine Comedy serving as the first controlled case study.

References
  1. Alighieri, Dante (1966-1967): La Commedia secondo l’antica vulgata. Ed. Giorgio Petrocchi. Milano: Mondadori.
  2. Boleda, Gemma (2020): “Distributional Semantics and Linguistic Theory”, in: Annual Review of Linguistics 6: 213-234.
  3. Erk, Katrin (2012): “Vector Space Models of Word Meaning and Phrase Meaning: A Survey”, in: Language and Linguistics Compass 6, 10: 635-653.
  4. Jockers, Matthew L. (2013): Macroanalysis: Digital Methods and Literary History. Urbana: University of Illinois Press.
  5. Navarro-Colorado, Borja (2018): “On Poetic Topic Modeling: Extracting Themes and Motifs From a Corpus of Spanish Poetry”, in: Frontiers in Digital Humanities 5: 15. DOI: 10.3389/fdigh.2018.00015.
  6. Pertile, Lino (1991): “L’antica fiamma: La metamorfosi del fuoco nella Commedia di Dante”, in: The Italianist 11, 1: 29-60. DOI: 10.1179/ita.1991.11.1.29.
  7. Reimers, Nils / Gurevych, Iryna (2019): “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks”, in: Proceedings of EMNLP-IJCNLP 2019: 3982-3992. DOI: 10.18653/v1/D19-1410.
  8. Song, Ruoci (2025): “‘Segni de l’antica fiamma’: Mapping Intertextuality between Virgil and Dante in the Divine Comedy”, in: Annali d’italianistica 43: 161-190. DOI: 10.65266/irst1090.
  9. Underwood, Ted (2019): Distant Horizons: Digital Evidence and Literary Change. Chicago: University of Chicago Press.
  10. van Zundert, Joris J. / Koolen, Marijn / Neugarten, Julia / Boot, Peter / van Hage, Willem / Mussmann, Ole (2022): “What Do We Talk About When We Talk About Topic?”, in: Proceedings of the Computational Humanities Research Conference 2022. CEUR Workshop Proceedings 3290: 398-410.
  11. Wang, Liang et al. (2024): “Multilingual E5 Text Embeddings: A Technical Report”. arXiv:2402.05672.