DH 2026

Daejeon, July 27–31

Fri, July 3114:00–15:30S024107
Long Paper

Towards a linked open data Elamite dictionary

Filippo Pedron
University of Helsinki · pedronfilippo05@gmail.com
Katrien De Graef
Ghent University · katrien.degraef@ugent.be
Shai Gordin
Ariel University · shaigo@ariel.ac.il
Timo Homburg
Mainz University of Applied Sciences · timo.homburg@gmx.de

The Elamite language, used in ancient southwestern Iran, remains one of the least digitally documented languages of the Ancient Near East. While other cuneiform-based languages, such as Sumerian, are increasingly accessible through online corpora, dictionaries, and digital tools, the digital documentation of Elamite is still in its infancy. This is striking, given the chronological depth and cultural significance of the Elamite textual tradition, spanning roughly three millennia and comprising almost 6,000 known texts from different periods and genres.

Since the end of the nineteenth century, Elamite studies have improved through the publication of ancient texts (e.g. Cameron 1948, Aliyari Babolaghni 2015), along with their grammatical and philological description (e.g. Paper 1955, Briceño Villalobos 2022). Yet no comprehensive grammar, dictionary, or corpus currently exists. Hinz and Koch’s 1987 lexical collection remains a fundamental reference work, but it is not a dictionary in the strict sense: entries are organised by spelling rather than lemma, ordering principles are inconsistent, and the work needs updating in light of later linguistic research and newly published or republished texts.

Digital resources remain equally limited. The most substantial digitisation effort has been the PFA Project in OCHRE, which provides access to the largest existing Elamite corpus, but only for one chronological and archival context (518/17 – 487 BCE). Moreover, the platform’s limited accessibility, the absence of freely downloadable and reusable data, and unstable searches have restricted its impact on wider Elamite studies (see Schloen and Prosser 2023). As a result, a rich textual tradition remains difficult to access, search, and analyse systematically.

Elamite studies are often included within Assyriology because of their geographical and chronological overlap with Mesopotamia and their use of cuneiform script, yet Elamite culture extends beyond the conventional boundaries of Mesopotamian studies. The first writing system attested in Elam, Proto-Elamite, belongs among the earliest writing systems in the world, and Elamite itself has not yet been securely attributed to any linguistic family. It therefore offers important opportunities for typological, comparative, and historical-linguistic inquiry.

Our project responds to this situation by developing new digital tools for the study of Elamite language and texts. Its aim is twofold: first, to provide more accessible, structured, and reusable resources for specialists; and second, to make Elamite materials available to a wider scholarly audience, thereby encouraging new research questions and new generations of scholars to engage with this important but understudied linguistic and cultural tradition.

The first major output is a digital lemma base for Elamite, conceived as a first step towards future linked open data dictionaries based on the OntoLex-Lemon paradigm. The primary input method was a set of shared tables in a Google Sheet environment. Given the fluid and often uncertain state of Elamite lexical knowledge, the annotations were created collaboratively by scholars and students of the Ancient Near East, guided by shared annotation guidelines that function as a living document (https://shorturl.at/GQUab).

The current lemma base, available on GitHub (https://digitalpasts.github.io/ALP-MEGA2024/), contains more than 15,000 lexemes, including approximately 4,300 nouns, 1,700 verbs, and more than 6,000 proper nouns. It is based on the XML version of Hinz and Koch’s lexical collection created by E. J. Smith, but the inherited material has been substantially reworked. Entries have been morphologically described, corrected, updated where appropriate, and enriched with machine-readable sense definitions. Loanwords from Sumerian (see also Chiarcos et al. 2018), Akkadian, and Old Persian have been identified and, where possible, linked to corresponding Wikidata Lexeme entries (Vrandečić 2012, Nielsen 2020). In this way, the lemma base transforms an older lexical collection into a structured, extensible, and digitally interoperable resource.

The second output is the use of this lemma base as an annotation resource for selected Elamite texts. These annotations are created manually according to the Web Annotation Data Model, using recogito-js (https://github.com/recogito/recogito-js). The annotation files are stored on GitHub and can be accessed in the browser, allowing individual textual forms to be connected with lexical, morphological, and semantic information in a transparent and reusable way with version control built-in. Thirdly, the project presents the first fonts specifically developed for Elamite, enabling a more accurate representation of Elamite transliteration and textual features in Unicode-based environments. Together, these outputs provide an initial digital infrastructure for Elamite philology and a basis for future community-driven annotation, correction, and expansion.

Figure 1: The Lemma base used to manually annotate a given Elamite text and display the text using newly developed fonts.

The usefulness of the lemma base is demonstrated through two case studies: one on Elamite adjectives and one on graphemic classifiers.

The first case study addresses the problem of parts of speech (POS) in Elamite. Since Elamite grammar remains insufficiently understood, POS are often divided broadly into nouns and verbs. Adjectives are frequently listed as a sub-class of nouns, since they may take similar suffixes and there is no clear suffix that consistently derives adjectives from nominal roots. Previous studies also added that class suffixes, such as -r and -p, mark adjectives in older periods, while possessive suffixes, such as -ni and -na, become more visible in later periods. Past/passive participles are also recognised as capable of functioning adjectivally.

In Smith’s XML version of Hinz and Koch, POS tags were automatically assigned, resulting in more than 500 entries classified as adjectives. Because this classification is partly mechanical and partly dependent on inherited assumptions, our team manually reviewed these entries in context and identified 196 roots actually used to qualify nouns. We then added fields to the lemma base to mark the suffixes attached to each form, allowing the data to be extracted, compared, and analysed diachronically.

The results confirm that suffixed nominal roots and participles are used adjectivally, but also refine earlier assumptions. Later periods show a marked increase in bare roots functioning as adjectives. In Achaemenid Elamite, where Iranian influence and loanwords may partly explain the phenomenon, bare roots make up almost 40% of all adjectival forms. Yet the development begins earlier: in Neo-Elamite, bare roots already account for approximately 14%. The analysis also shows that there is no distinct adjectival base in Elamite, since the same roots may take different suffixes when used adjectivally. Finally, contrary to previous assumptions, class suffixes remain relatively stable throughout the attested periods, while possessive suffixes peak in Neo-Elamite and decline in Achaemenid Elamite, probably due to the increasing use of bare roots.

Figure 2: normalized heatmap showing the means of creating adjectives throughout different periods. 

The second case study concerns graphemic classifiers. In cuneiform writing systems, classifiers signal semantic categories of words, usually through preposed or postposed signs. Elamite employs a much smaller repertoire than Sumerian. Some classifiers correspond directly to Mesopotamian usage, such as DINGIR for divine names, MUNUS for female personal names or female-related lexemes, and GIŠ for wooden objects. Others reflect specifically Elamite developments, such as AŠ, marking place names, and final MEŠ, originally the Sumerian plural marker, used in Elamite to indicate a preceding logogram or pseudo-logogram.

Despite their importance for understanding the relation between writing, language, and semantic categorisation, the diachronic development of classifiers in Elamite has not yet been systematically studied. The lemma base now makes such an analysis possible. In an exploratory study, the classifier notations used by Hinz and Koch were extracted automatically from the transliteration column and mapped onto classifier_front and classifier_end fields. To enable diachronic comparison, the heterogeneous period labels of Hinz and Koch were consolidated into four broad chronological phases: Old Elamite, Middle Elamite, Neo-Elamite, and Late/Achaemenid Elamite.

Across nouns and proper names, classifier usage increases steadily over time, with a particularly strong rise in the late period. Geographical names shift from postposed KI in the old period to preposed AŠ from the middle period onward, with a brief appearance of URU in the Neo-Elamite period, probably reflecting Neo-Assyrian influence. Personal names display a parallel development: DIŠ and MUNUS dominate the old and middle periods, BE appears in the Neo-Elamite period, and HAL becomes the principal classifier in the late period. For nouns, the picture is more complex: MEŠ appears from the middle period onward, BE enters in the Neo-Elamite period, and HAL increases steadily from Neo-Elamite into Late/Achaemenid Elamite. Overall, these patterns reveal both a quantitative increase in classifier usage and a qualitative shift from inherited Mesopotamian sortal classifiers to originally mensural classifiers, such as AŠ and HAL, refunctionalised as sortal markers within the Elamite tradition.

These case studies demonstrate the potential of the Elamite lemma base to reassess long-standing linguistic questions through more structured and transparent datasets. The project shows how digital tools can support both the confirmation and revision of previous scholarship, while creating new possibilities for linguistic and cultural-historical analysis. By publishing the lemma base, annotations, and fonts openly, we aim to provide a reusable infrastructure for Elamite studies, involve the existing community of specialists, and spark renewed interest among new generations of scholars. The project therefore offers not only technical advances, but also a model for how a small and understudied field can be strengthened through open, collaborative, and community-oriented digital philology.

References
  1. Aliyari Babolghani, S. (2015): The Elamite Version of Darius the Great’s Inscription at Bisotun. Tehran: Nashr-e Markaz (in Persian).
  2. Arfaee, A. (2008): Persepolis Fortification Tablets: Fort. and Teh. Texts. Tehran.
  3. Bavant, M. (2019): “À propos d'éventuels cognats caucasiques en élamite”, in: Bulletin de la Société de Linguistique de Paris 114, 1: 341–383.
  4. Bosque-Gil, Julia / Gracia, Jorge / McCrae, John / Cimiano, Philipp / Stolk, Sander / Khan, Fahad / Depuydt, Katrien / de Does, Jesse / Frontini, Francesca / Kernerman, Ilan (2019): “The OntoLex Lemon lexicography module”, in: W3C Community Report.
  5. Briceño Villalobos, J. E. (2022): “Achaemenid Elamite and Old Persian Indefinites: A Comparative View”, in: Bianconi, M. et al. (eds.): Ancient Indo-European Languages between Linguistics and Philology. Leiden: Brill: 48–87.
  6. Cameron, G. G. (1948): Persepolis Treasury Tablets. Oriental Institute Publications 65. Chicago, IL: The University of Chicago Press.
  7. Chiarcos, C. / Pagé-Perron, É. / Khait, I. / Schenk, N. / Reckling, L. (2018): “Towards a Linked Open Data Edition of Sumerian Corpora”, in: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). Miyazaki: European Language Resources Association (ELRA).
  8. Grillot-Susini, F. (1987): Éléments de grammaire élamite. Paris: Éditions Recherche sur les Civilisations.
  9. Grillot-Susini, F. (2008): L’Élamite: Éléments de grammaire. Paris: Geuthner.
  10. Hallock, R. T. (1969): The Persepolis Fortification Tablets. Chicago: The University of Chicago Press.
  11. Henkelman, W. F. M. (2022): “God is in the detail: The divine determinative and the expression of animacy in Elamite, with an appendix on the Achaemenid calendar”, in: Cancik-Kirschbaum, E. / Schrakamp, I. (eds.): Transfer, Adaptation und Neukonfiguration von Schrift- und Sprachwissen im Alten Orient. Wiesbaden: Harrassowitz Verlag: 405–477.
  12. Hinz, W. / Koch, H. (1987): Elamisches Wörterbuch. Berlin: Reimer.
  13. Khačikjan, M. L. (1998): The Elamite Language. Documenta Asiana 4. Rome.
  14. König, W. F. (1965): Die elamischen Königsinschriften, in: Archiv für Orientforschung 16. Graz.
  15. Labat, R. (1957): “Structure de la langue élamite”, in: Conférences de l’Institut de Linguistique de l’Université de Paris 10: 23–42.
  16. Malbran-Labat, F. (1995): Les inscriptions royales de Suse. Paris: Réunion des Musées Nationaux.
  17. McAlpin, D. W. (1981): Proto-Elamo-Dravidian: The Evidence and its Implications. Philadelphia: The American Philosophical Society.
  18. Nielsen, Finn (2020): “Lexemes in Wikidata: 2020 status”, in: Proceedings of the 7th Workshop on Linked Data in Linguistics (LDL-2020). Marseille: European Language Resources Association: 82–86.
  19. Paper, H. H. (1955): The Phonology and Morphology of Royal Achaemenid Elamite. Ann Arbor: The University of Michigan Press.
  20. Pedron, F. (2023): Elamite and Dravidian: A Reassessment. PB, Demmy 1/8.
  21. Reiner, E. (1969): “The Elamite Language”, in: Friedrich, J. / Reiner, E. / Kammenhuber, A. / Neumann, G. / Heubeck, A. (eds.): Altkleinasiatische Sprachen. Leiden.
  22. Saber, A. P. (2017): A New Edition of the Elamite Version of the Behistun Inscription (I). Karaj, Iran.
  23. Sanderson, Robert / Ciccarese, Paolo / Van de Sompel, Herbert (2013): “Designing the W3C open annotation data model”, in: Proceedings of the 5th Annual ACM Web Science Conference (WebSci '13). New York, NY: Association for Computing Machinery: 366–375. DOI: 10.1145/2464464.2464474.
  24. Schloen, S. R. / Prosser, M. C. (2023): Database Computing for Scholarly Research: Case Studies Using the Online Cultural and Historical Research Environment. 1st ed. Quantitative Methods in the Humanities and Social Sciences. Cham: Springer International Publishing.
  25. Selz, G. J. / Grinevald, C. / Goldwasser, O. (2017): “The Question of Sumerian ‘Determinatives’: Inventory, Classifier Analysis, and Comparison to Egyptian Classifiers from the Linguistic Perspective of Noun Classifiers”, in: Werning, D. A. (ed.): Proceedings of the Fifth Conference on Egyptian-Coptic Linguistics (Crossroads V), Berlin February 17–20, 2016, in: Lingua Aegyptia 25: 281–344.
  26. Smith, E. J. (2007): “Phonological reconstruction of a dead language using the gradual learning algorithm”, in: Proceedings of the Ninth Meeting of the ACL Special Interest Group in Computational Morphology and Phonology: 57–64.
  27. Stolper, M. W. (1984): Texts from Tall-i Malyan I: Elamite Administrative Texts. Philadelphia.
  28. Stolper, M. W. (2004): “Elamite”, in: Woodard, R. D. (ed.): The Cambridge Encyclopedia of the World’s Ancient Languages. Cambridge: Cambridge University Press: 60–94.
  29. Tavernier, J. (2018): “The Elamite language”, in: Alvarez-Mon, J. / Basello, G. P. / Wicks, Y. (eds.): The Elamite World. London: Routledge: 416–450.
  30. Vrandečić, Denny (2012): “Wikidata: a new platform for collaborative data collection”, in: Proceedings of the 21st International Conference on World Wide Web (WWW '12 Companion). New York, NY: Association for Computing Machinery: 1063–1064. DOI: 10.1145/2187980.2188242.
  31. Wagensonner, K. (2021): “Classifiers between Euphrates and Tigris: On development and use of noun categorization in cuneiform script”, in: Gabriel, G. / Overmann, K. A. / Payne, A. (eds.): Signs – Sounds – Semantics: Nature and Transformation of Writing Systems in the Ancient Near East. Wiener Offene Orientalistik 13. Wien: Ugarit-Verlag: 171–212.