DH 2026

Daejeon, July 27–31

Fri, July 3109:00–10:30S049206-208
Short Paper

«Le varianti della rosa». From annotation to engagement: storytelling as a model for digital scholarly editions using DSL, XML-TEI, and LLMs

Christian D'Agata
University of Catania, Italy · christian.dagata@gmail.com

Introduction: Umberto Eco and «Le varianti della rosa»

Umberto Eco is one of the leading Italian intellectuals, semioticians, and writers of the late twentieth century

For a bibliography updated to the early 2000s, see Contursi (2005). For a recent bibliography, see the bibliography in Palazzolo (2019) and D'Agata (2025).

, internationally renowned for his 1980 masterpiece The Name of the Rose, set in a medieval monastery. Less well known is that, in the final years of his career, Eco promoted a thorough revision of the novel, producing a fully revised and corrected edition published in 2012

The first edition was published by Bompiani, Milan, 1980 (Eco 1980/1983); while the revised and corrected edition was published by Bompiani, Milan, 2012 (Eco 2012). The edition currently on sale is based on the latter, published by La Nave di Teseo with illustrations by Eco himself (La Nave di Teseo, Milan, 2020).

. This paper presents, as a case study, the portal Le varianti della rosa

www.variantidellarosa.it

, developed since 2020 and dedicated to the variants of The Name of the Rose through an approach that combines philology and hermeneutics within a fully digital methodology. In the context of a broader rethinking of textuality – where digital scholarly editions too often address only academic audiences – the project proposes a digital edition conceived as a hub of multiple experiences (an ecosystem integrating lexicography [Savoca 2000], philology, education, and multimedia) aimed at a wider public. Central to this proposal is the concept of engagement, articulated through the IDEA paradigm (Interpretation, Didactics, Edition, Annotation) developed over the course of the project and expressed through the notion of hyperedition, already implemented in portals such as Pirandello Nazionale (Subialka / Di Silvestro / Sichera 2019; D’Agata / Di Silvestro / Sichera 2022), Paves-e (D’Agata et al. 2024), and Verismo Digitale (Barbarino et al. 2024). Building on this framework, the project now moves toward an expanded model – IDEAS – where Storytelling becomes a structural component of the digital scholarly edition, bridging critical rigor and public accessibility.

From IDEA to IDEA(S): Hyperedition and Storytelling

The ‘IDEA’ paradigm emerges from the need to overcome the traditional compartmentalization of digital philology (D’Agata 2022). Interpretation refers to the set of exegetical and hermeneutic practices (Sichera 2017) accompanying a text and which, within the hyperedition, are no longer confined to footnotes. Instead, they unfold through multiple layers of annotation and visualization, involving an active role of both encoder and reader, who select what to mark and display according to their specific textual focus. Didactics concerns the potential of transforming the portal into a tool for schools and universities, through interactive exercises, guided pathways, and multimedia materials that turn textual complexity into an opportunity for continuous discovery. Edition identifies the core processes of textual representation based on contemporary digital philology, while Annotation involves the use of formal languages (XML-TEI and DSL) that make the text (and its commentary) computable, queryable, and reusable (Wilkinson, et al 2016) – a process that, in turn, cyclically depends on the critical interpretation of the editor.

The shift from IDEA to IDEA(S) introduces Storytelling as a fifth, integrative dimension (Fig. 1). In this model, storytelling is not conceived as a simplification or popularization of scholarly content, but as a strategy for structuring and communicating philological knowledge. Variants are presented as narrative moments in the life of the text, allowing readers to follow the evolution of authorial choices, stylistic adjustments, and semantic shifts over time.

Figure The IDEA(S) paradigm

More specifically, the digital edition of the variants was developed through the collation of the 1980 and 2012 texts (acquired via OCR) using automatic collation algorithms (CollateX

https://collatex.net

, Juxta

Cfr. https://journalofdigitalhumanities.org/3-1/juxta-commons/. As of today (05/2026) the Juxta commons website is unavailable.

, and VarianceViewer

https://github.com/cs6-uniwue/Variance-Viewer

). The results were compared and manually revised, ultimately producing XML- TEI

Cfr. https://tei-c.org

encoding of the variants (<app>, <lem>, <rdg>), which was displayed using EVT

http://evt.labcd.unipi.it

(Fig, 2).

Figure Digital Scholarly Edition visualized using EVT software

The variants were then annotated with a Domain Specific Language (Parr 2007; Parr 2014; Zenzaro, et al 2025) through the Euporia software (Boschetti, et al. 2023; Boschetti / Del Grosso 2020; Mugelli et al. 2016), developed by CNR-ILC. The DSL provided generic annotations concerning the communicative context of each variant (the characters involved, the thematic and narrative elements of the paragraph) as well as linguistic annotations (Fig. 3), particularly based on Tullio De Mauro’s (1999-2000) usage labels (“fundamental vocabulary,” “high usage,” “low usage,” “literary,” “specialized”). This framework enabled a range of statistical analyses, which helped define Eco’s corrective interventions not as simplification, but as a rhapsodic exercise in labor limae.

Figure Example of annotation in Euporia (2020)

Engagement and Road Map

From the standpoint of engagement, the crucial elements lie not only in the richness of philological data, but also in their transformation into narrative experiences. The portal includes a section dedicated to the storytelling of variants, where significant transformations are visualized using TRAViz

http://www.traviz.vizcovery.org

, conceptual maps, infographics, and thematic pathways (Fig. 4).

Figure Storytelling with TRAViz and infographics

Another example of storytelling is with StoymapJS and Timeline JS (Fig. 5). The aim is to make the variant not a technical artifact accessible only to specialists, but an entry point into understanding the novel’s construction of meaning

A new release planned for 2026 consolidates the IDEA(S) paradigm by integrating research infrastructure and storytelling-driven engagement: 1)A variant search engine, developed as a web application based on eXist-db, enabling complex queries on words, usage labels, characters, and themes, as represented in the Euporia-generated annotations; 2) An experimental linguistic and interpretive commentary system, generated in natural language and supported by large language models (GPT, Gemini, Claude, DeepSeek), combining philological accuracy with narrative-driven strategies to present variants as evolving stories of textual change, making scholarly interpretation accessible to non-specialist audiences without sacrificing critical depth.

.

Figure StorymapJS and TimelineJS (Knight lab)

Conclusion. An example of IDEA(S): variant XIII

To illustrate the IDEA(S) model in practice, we conclude with a single variant treated across all the encoding layers it traverses in the portal. The case is variant XIII of The Name of the Rose, located at the opening of the description of Guglielmo di Baskerville. In NR80/83, William’s portrait begins with a sentence overtly modeled on the incipit of A. Conan Doyle’s A Study in Scarlet: «Era dunque l’apparenza fisica di frate Guglielmo tale da attirare l’attenzione dell'osservatore più distratto» – a near-translation of «His very person and appearance were such as to strike the attention of the most casual observer». In NR12, Eco silently removes the Holmesian opening and rewrites the paragraph so that the description now starts with the more neutral «La statura di frate Guglielmo».

Layer 1: XML-TEI encoding

At the philological-data layer, the variant is encoded according to the TEI parallel-segmentation model:

Figure XML-TEI encoding and query via eXist-db app

This level guarantees machine-readability and visualization through eXist-db (Fig. 6); it makes the variant queryable, FAIR-compliant, and reusable for subsequent computational analysis.

Layer 2: DSL annotation

At the interpretive-annotation layer, the same variant is encoded in the Euporia DSL, which captures dimensions that TEI cannot easily express – speaker, topic, hermeneutic cause, possible-world status:

& XIII
*[0] @narratore §Guglielmo/descrizione
*[1] {Era dunque l'apparenza fisica di frate Guglielmo tale da attirare l'attenzione dell'osservatore più distratto. } = §alleggerimento/maior §riferimentoSherlock §variazionePersonaggio §mondoPossibile/globale
*[2] {La sua statura} : <La statura di frate Guglielmo> = §esplicitazione

Layer 3: Natural-language storytelling

At the public-facing layer, the same variant becomes a short narrative addressed to a non-specialist reader:

The Holmesian opening. NR80 begins its portrait of Guglielmo with a phrase clearly derived from Conan Doyle: “It was, therefore, the physical appearance of Brother Guglielmo that was such as to attract the attention of even the most casual observer.” The echo is that of the first description of Sherlock Holmes in A Study in Scarlet (“His very person and appearance were such as to strike the attention of the most casual observer”). NR12 erases everything. Thus begins the systematic “de-Holmesification” of the character, which continues in & XIV and XV...

This third layer, drafted by the editor and refined through LLM-assisted commentary (with prompts conditioned on the DSL annotation as structured input)

A related reflection on the interaction between DSLs and LLMs has been accepted for presentation at the AIUCD 2026 conference: D'Agata, Christian / Del Grosso, Angelo M. / Boschetti, Federico, "Per una formazione del filologo computazionale del futuro: Domain Specific Language e Large Language Model" (forthcoming).

, is what the IDEA(S) model adds to the digital scholarly edition: a narrative register in which the data of layer 1 and the interpretation of layer 2 become legible to a non-academic audience without losing critical accuracy. The same DSL annotation also drives a visual extension of the storytelling layer: LLM like Claude generates HTML cards that render the variant as an interactive scene (NR83 vs NR12 typographic blocks, drop caps, collapsible philological apparatus), while Midjourney produces image illustrations (Fig. 7).

Figure Annotation experiment developed using DSL and LLM

The model in practice

Variant XIII demonstrates the architecture of the hyperedition we propose. The same unit of textual change is rendered as TEI-data, as DSL-interpretation, and as narrative-experience, with traceable links between the three. The Storytelling layer does not replace the philological apparatus: it is generated from it. In this sense, IDEA(S) is not an additive paradigm: Storytelling is what holds together Interpretation, Didactics, Edition, and Annotation when the digital scholarly edition steps outside the academy. Le varianti della rosa offers a working prototype of this replicable workflow – collation, TEI encoding, DSL annotation, narrative rendering with LLMs – for digital editions that aim simultaneously at scholarly rigor and public engagement.

References
  1. Barbarino, Liborio P. / Conti, Elisa / D'Agata, Christian / Grasso, Miryam / Martines Ninna M. L. / Vitale, Eliana (2024). “Verismo Digitale. Per un’edizione digitale commentata delle opere di Verga, Capuana, De Roberto”, in: Me.Te. Digitali. Mediterraneo in rete tra testi e contesti, Proceedings del XIII Convegno Annuale AIUCD, Catania 28-30 maggio 2024, Università di Catania, edited by A. Di Silvestro, D. Spampinato, Catania, AIUCD: 252-259.
  2. Boschetti, Federico / Del Grosso, Angelo M. (2020). “L’annotazione di testi storico-letterari al tempo dei social media”, in: Italica Wratislaviensia 11 (1): 65–99.
  3. Boschetti, Federico / Bambaci, Luigi / Del Grosso, Angelo M. / Mugelli, Gloria / Khan, Fahad / Bellandi, Andrea / Taddei, Andrea (2023). “Collaborative and Multidisciplinary Annotations of Ancient Texts: The Euporia System”, in: The Ancient World Goes Digital: Case Studies on Archaeology, Texts, Online Publishing, Digital Archiving, and Preservation, edited by V.B Juloux, et al., Leiden, Brill Academic Publishers: 172-223.
  4. Contursi, James L. (2005). Umberto Eco: An Annotated Bibliography of First and Important Editions, Minneapolis, Minnesota Bookman Publications.
  5. D’Agata, Christian (2022). “L'edizione scientifica digitale estesa de "Il nome della rosa": modellizzazione, workflow e il paradigma IDEA”, in: Umanistica Digitale, 14.
  6. D’Agata, Christian (2025). Nel nome della rosa. Apocalissi, genesi, varianti, Strasbourg, Eliphi.
  7. D’Agata, Christian / Di Silvestro, Antonio / Sichera, Antonio (2022). “Edizione critica, edizione digitale, Hyperedizione. «Il fu Mattia Pascal» come paradigma dell’edizione digitale dell’Opera Omnia di Luigi Pirandello”, in: Bollettino - Centro di studi filologici e linguistici siciliani, 33.
  8. D’Agata, Christian / Del Grosso, Angelo M. / Nay, Laura / Palazzolo, Giuseppe / Sichera, Antonio / Spampinato, Daria (2024). “Paves-e: per una hyperedizione dell’opera di Cesare Pavese”, in: Me.Te. Digitali. Mediterraneo in rete tra testi e contesti, Proceedings del XIII Convegno Annuale AIUCD, Catania 28-30 maggio 2024, Università di Catania, edited by A. Di Silvestro, D. Spampinato, Catania, AIUCD: 191-196.
  9. De Mauro, Tullio (1999-2000) Grande dizionario italiano dell’uso, Torino, UTET, 1999-2000, 6 voll. Con DVD-ROM; vol. 7, Nuove parole italiane dell’uso, 2003, con DVD-ROM; vol. 8, Nuove parole italiane dell’uso II, 2007, con penna USB.
  10. Eco, Umberto (1980/1983). Il nome della rosa, Milano, Bompiani.
  11. Eco, Umberto (2012). Il nome della rosa, I edizione riveduta e corretta, Milano, Bompiani.
  12. Mugelli, Gloria / Boschetti, Federico / Del Gratta, Riccardo / Del Grosso, Angelo M. / Khan, Fahad / Taddei, Andrea (2016), “A User-Centered Design to Annotate Ritual Facts in Ancient Greek Tragedies”, in: Bulletin of the Institute of Classical Studies, 59 (2): 103-120.
  13. Palazzolo, Giuseppe (2019). Umberto Eco. Epifanie, ossessioni, gnosi, Lentini (SR), Duetredue.
  14. Parr, Terence (2007). The Definitive ANTLR Reference: Building Domain-Specific Languages, Raleigh (NC), Pragmatic Bookshelf.
  15. Parr, Terence (2014). Language Implementation Patterns Create Your Own Domain-Specific and General Programming Languages, Raleigh (NC), Pragmatic Bookshelf.
  16. Savoca, Giuseppe (2000). Lessicografia letteraria e metodo concordanziale. Firenze, Olschki.
  17. Sichera, Antonio (2017). Ermeneutiche. Punti di vista sul confine. Leonforte (En), Euno Edizioni.
  18. Subialka, Michael / Di Silvestro, Antonio / Sichera, Antonio (2019). “The future of Pirandello: On the New Digital Edition of Pirandello's Opera omnia. Antonio Sichera and Antonio Di Silvestro in conversation with Michael Subialka”, in: The Journal of The Pirandello Society of America, XXXII: 107-16.
  19. Wilkinson, Mark D. / et al. (2016). “The FAIR Guiding Principles for scientific data management and stewardship”, in: Sci Data 3, 160018. https://doi.org/10.1038/sdata.2016.18.
  20. Zenzaro, Simone / Boschetti, Federico / Del Grosso, Angelo M. (2025). “Making digital scholarly editions based on Domain Specific Languages”, in: Digital editing and Publishing in the Twenty-first Century, edited by J. O'Sullivan et al., Edinburgh, Scottish Universities Press: 141-164.