DH 2026

Daejeon, July 27–31

Thu, July 3013:40–15:10S057209-211
Short Paper

DISSILEX: A machine-operable lexico-semantic network and valency lexicon for medieval Latin

Gideon Kotzé
Masaryk University, Czech Republic · gideon.kotze@mail.muni.cz
Robert L.J. Shaw
Masaryk University, Czech Republic · robert.shaw@mail.muni.cz
Katia Riccardo
Masaryk University, Czech Republic · katia.riccardo@mail.muni.cz
Katalin Suba
Masaryk University, Czech Republic · katalin.suba@mail.muni.cz
Davor Salihović
University of Antwerp, Belgium · davor.salihovic@uantwerpen.be
David Zbíral
Masaryk University, Czech Republic · david.zbiral@mail.muni.cz

We present DISSILEX, a new lexicographic resource of verbs and verbal expressions in medieval Latin, featuring a detailed valency lexicon and connections to a large lexico-semantic network of Latin and modern-English concepts, defining and contextualizing the verb’s meaning and argument structure. DISSILEX is published as a Streamlit web application, as well as a versioned JSON and a SQLite snapshot via Zenodo. The resource is rooted in the morphosyntactic-semantic annotation of medieval inquisition records for the purposes of historical research. As of May 2026, DISSILEX contains 909 Latin Actions (verbs and verbal expressions, such as do penitentiam, “to assign penance”) and 788 linked Concepts (mostly nouns, verbs (gerunds) and adjectives, including multi-word expressions), connected by 2,476 labelled relations across 13 edge types.

DISSILEX is designed for integration into the Linguistic Linked Open Data (LLOD) ecosystem (Chiarcos et al. 2013; Cimiano et al. 2020) through the Linking Latin (LiLa) knowledge base (Passarotti et al. 2020). IDs from the LiLa Lemma Bank (Mambrini and Passarotti 2023) link DISSILEX to Latin WordNet (Minozzi 2017), corpora, dictionaries, valencies, ontologies, and NLP tools. It also links to Princeton WordNet (PWN; Miller 1995) via the Collaborative Inter-Lingual Index (Bond et al., 2016). 97.6% of single-word Actions reference LiLa, while around 80% link to a WordNet synset (PWN 3.0 and/or 3.1); 134 single-word Actions lack a WordNet equivalent. This dual alignment layer of form (Lemma Bank) and meaning (WordNet) enables LLOD integration.

Rooted in the domain of inquisitorial records, DISSILEX covers general and more subject-specific meanings, with both standard (synonym, hypernym, etc.) and less canonical relations. DISSILEX is a product of Computer-Assisted Semantic Text Modelling (CASTEMO; Zbíral et al. 2025), an approach to modeling statements as a quadruple-structure of subject(s), predicate(s) and two objects, creating a thickly connected network of data points stored in a Neo4j graph database.

Each Action represents a distinct meaning with an appropriate valency structure. Actant slots (subject, object 1, object 2) may have three different valency types:

  • entity type valency: which entity types are allowed in the given actant slot, such as Person or Group.
  • morphosyntactic valency: applies a formalized notation to indicate morphosyntactic constraints. For example: acc. | acc. with inf. | "quod" for accusative, accusative-with-infinitive, or a “quod”-clause.
  • semantic valency: names the actant’s semantic role in English, such as “baptized person”.

Our valency component displays similarities to VerbNet (Schuler 2005), the Valency Lexicon of Czech Verbs (Žabokrtský 2005) and Latin VALLEX (Passarotti et al. 2016); similarly, there are parallels between our semantic network and WordNet, as well as the theoretical foundations of models like Role and Reference Grammar (Van Valin / Foley 1980). The multilingual Latin-English coverage also places DISSILEX in the tradition of BabelNet (Navigli/Ponzetto 2012).

We illustrate the model with the example Action baptizo (“to baptize”). The subject slot requires “person” or “group” in the nominative and the semantic role of “baptizer”. “person” or “group” is required in the accusative, with the semantic role of recipient of baptism. The passive counterpart baptizatus/a, a separate Action, links to baptizo with the Subject/Actant 1 Reciprocal (SAR) type, a semantic relation that captures pairs where the subject and first object roles are swapped. Both baptizo and baptizatus/a in turn link to the Concept baptism via Action/Event Equivalent (AEE), while baptism links as synonym (SYN) to the Latin baptismus. Work on explicit active/passive meaning attributes is partially implemented.

Lemmas include multi-word expressions (e.g., do penitentiam, "give penance"), enabling deeper taxonomy paths such as sanctioninstitution of social controlinstitutionentity.

Aside from the aforementioned AEE and SAR, we also introduce these domain-motivated relation types:

  • Property Reciprocal (PRR) captures reciprocal bidirectional relationships (e.g. “parent” and “child”).
  • The Actant Semantics relations (SUS, A1S, A2S) assign FrameNet-style semantic roles to actant slots (e.g. loquor; “to speak, to sb - about st/sb”), to the appropriate semantic role concept (speaker, interlocutor, topic).

Alignment to external resources is based on exact matching, taking definitions into account. If a DISSILEX sense is more specific than a WordNet synset, we create an Action corresponding to the WordNet meaning, which then forms the hypernym of this more specific Action. Synonymic Actions are linked and need to share the same hypernym Action.

Figure 1. Here we utilize the running example Action baptizo ("to baptize") to demonstrate our model. Each Action has three valency types in each actant slot. Relations are represented as edges. Domain edges (SAR, AEE, SUS, A1S, SCL (superclass, our term for hypernym), and SYN) are shown alongside external relations (LiLa, PWN 3.0/3.1, Latin WordNet via shared PWN 3.1 offsets, CILI index).

Actions for which we have not found an appropriate WordNet meaning are marked explicitly in the source dataset, making the coverage gap machine-readable. For example, one meaning of consolo refers to the sense of performing a specific form of unorthodox baptismal ritual. Such domain-specific markers would allow historical scholars to identify Actions that are specific to Inquisition-related materials, such as those involving inquisitors and deponents in subject/object positions. The growing knowledge graph enables a deeper understanding of the subject material, assisting with various forms of distant reading and quantitative analysis methods applied to both historical research and natural language understanding.

Quality control mechanisms combine guidelines, in-app validation at entry (such as required fields and valency type symmetry), validations during data compilation, and entity status transparently included in the public data (such as approved = cross-checked, meeting quality requirements and guidelines; pending: further checks are required).

Ongoing work includes alignment to the PreMOn ontology (Corcoglioniti et al. 2016), an extension of OntoLex-Lemon (Bosque-Gil et al. 2017) that represents semantic roles and predicate structures. This would represent our lemmas, relations and valencies via URIs in the LLOD cloud and would allow external discovery of DISSILEX data through SPARQL endpoints, thus contributing to improved computational access to lexicographic data of medieval Latin.

References
  1. Baker, Collin F. / Fillmore, Charles J. / Lowe, John B. (1998): "The Berkeley FrameNet project", in: Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics, Volume 1. Montreal: Association for Computational Linguistics, pp. 86-90. DOI: 10.3115/980845.980860.
  2. Bond, Francis / Vossen, Piek / McCrae, John P. / Fellbaum, Christiane (2016): "CILI: the Collaborative Interlingual Index", in: Fellbaum, Christiane / Vossen, Piek / Mititelu, Verginica Barbu / Forascu, Corina (eds.): Proceedings of the 8th Global WordNet Conference (GWC). Bucharest: Global Wordnet Association, 50–57. DOI: 10.18653/v1/2016.gwc-1.9.
  3. McCrae, John P. / Bosque-Gil, Julia / Gracia, Jorge / Buitelaar, Paul / Cimiano, Philipp (2017): "The Ontolex-Lemon model: development and applications", in: Proceedings of eLex 2017 conference. Leiden, 587-597.
  4. Chiarcos, Christian / McCrae, John P. / Cimiano, Philipp / Fellbaum, Christiane (2013): "Towards Open Data for Linguistics: Linguistic Linked Data", in: Oltramari, Alessandro / Vossen, Piek / Qin, Lu / Hovy, Eduard (eds.): New Trends of Research in Ontologies and Lexical Resources. Ideas, Projects, Systems. Berlin / Heidelberg: Springer, 7-25. DOI: 10.1007/978-3-642-31782-8_2.
  5. Cimiano, Philipp / Chiarcos, Christian / McCrae, John P. / Gracia, Jorge (2020): "Linguistic Linked Open Data Cloud", in: Cimiano, Philipp / Chiarcos, Christian / McCrae, John P. / Gracia, Jorge (eds.): Linguistic Linked Data: Representation, Generation and Applications. Cham: Springer International Publishing, 29-41. DOI: 10.1007/978-3-030-30225-2.
  6. Corcoglioniti, Francesco / Rospocher, Marco / Aprosio, Alessio Palmero / Tonelli, Sara (2016): "PreMOn: a Lemon Extension for Exposing Predicate Models as Linked Data", in: Calzolari, Nicoletta et al. (eds.): Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16). Portorož: European Language Resources Association (ELRA), 877-884. <https://aclanthology.org/L16-1141/>
  7. Mambrini, Francesco / Passarotti, Marco Carlo (2023): "The LiLa Lemma Bank: A Knowledge Base of Latin Canonical Forms", in: Journal of Open Humanities Data 9, 1. DOI: 10.5334/johd.145.
  8. Miller, George A. (1995): "WordNet: a lexical database for English", in: Communications of the ACM 38, 11: 39-41. DOI: 10.1145/219717.219748.
  9. Minozzi, Stefano (2017): "Latin WordNet, una rete di conoscenza semantica per il latino e alcune ipotesi di utilizzo nel campo dell’information retrieval", in: Mastandrea, Paolo (ed.): Strumenti digitali e collaborativi per le Scienze dell’Antichità. Venezia: Edizioni Ca' Foscari (Antichistica 14). DOI: 10.14277/6969-182-9/ANT-14-10.
  10. Navigli, Roberto / Ponzetto, Simone Paolo (2012): "BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network", in: Artificial Intelligence 193: 217-250. DOI: 10.1016/j.artint.2012.07.001.
  11. Passarotti, Marco / Mambrini, Francesco / Franzini, Greta / Cecchini, Flavio Massimiliano / Litta, Eleonora / Moretti, Giovanni / Ruffolo, Paolo / Sprugnoli, Rachele (2020): "Interlinking through Lemmas. The Lexical Collection of the LiLa Knowledge Base of Linguistic Resources for Latin", in: Studi e Saggi Linguistici 58, 1: 177-212. DOI: 10.4454/ssl.v58i1.277.
  12. Passarotti, Marco / Saavedra, Berta González / Onambele, Christophe (2016): "Latin Vallex. A Treebank-based Semantic Valency Lexicon for Latin", in: Calzolari, Nicoletta / Choukri, Khalid / Declerck, Thierry / Goggi, Sara / Grobelnik, Marko / Maegaard, Bente / Mariani, Joseph / Mazo, Helene / Moreno, Asuncion / Odijk, Jan / Piperidis, Stelios (eds.): Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16). Portorož: European Language Resources Association (ELRA), 2599-2606. <https://aclanthology.org/L16-1414/>
  13. Schuler, Karin Kipper (2005): VerbNet: A broad-coverage, comprehensive verb lexicon. Ph.D. thesis, University of Pennsylvania. <https://www.proquest.com/openview/7ca4b1b9093522a7d8089ff2e987e74e/>
  14. Van Valin Jr., Robert D. / Foley, William A. (1980): "Role and Reference Grammar", in: Moravcsik, Edith A. / Wirth, Jessica R. (eds.): Syntax and Semantics 13: Current Approaches to Syntax. New York: Academic Press, 329-352. DOI: 10.1163/9789004373105_014.
  15. Žabokrtský, Zdeněk (2005): Valency lexicon of Czech verbs. Ph.D. thesis, Charles University, Prague.
  16. Zbíral, David / Shaw, Robert L.J. / Hanák, Petr / Hampers, Tomáš / Mertel, Adam (2025): From texts to structured data: Building knowledge graphs through Computer-Assisted Semantic Text Modelling (CASTEMO). Brno: Masaryk University. <https://docs.religionistika.phil.muni.cz/books/from-texts-to-structured-data-building-knowledge-graphs-through-computer-assisted-semantic-text-modelling-castemo>