DH 2026

Daejeon, July 27–31

Fri, July 3114:00–15:30S068106
Short Paper

Narratives of maritime risk: a Natural Language Processing approach to reassess ancient seafaring strategies in the Mediterranean Sea

Ruhi Mahadeshwar
University of Groningen, The Netherlands · r.u.mahadeshwar@rug.nl
Manuela Ritondale
University of Groningen, The Netherlands · manuela.ritondale@gmail.com
Malvina Nissim
University of Groningen, The Netherlands · m.nissim@rug.nl

Introduction

Modelling approaches to ancient seafaring have expanded significantly in recent decades (see (Slayton et al. 2025) for a recent review), yet the role of risk perception in shaping navigational choices has not been formally addressed or integrated into such models (an exception being (Ritondale 2022)). Risk is widely recognised as a central factor in voyaging decisions. Still, it is often assumed that mariners across periods and cultures confronted and managed risk in essentially similar ways, and that “risk management and curative behaviour are unifying trends in all traditions of seamanship” (Maarleveld 2004: 142). Consequently, seafaring modelling approaches frequently treat the safest routes as the most probable ones. This assumption can prove to be simplistic, as it does not consider subjective perceptions and knowledge (e.g., Polybius, Histories, 1.39.1), cultural attitudes, and the reality of competing or shifting risk factors (e.g., Polybius, Histories, 1.54.1). Moreover, the prevailing narrative that ancient mariners consistently preferred coastal navigation, despite the fact that “coastal” sailing is itself a more complex and variable practice than the term suggests, does not align neatly with the more nuanced picture emerging from shipwreck evidence and the textual record (Ritondale 2022).

These observations raise several questions that will guide our study:

  • How were maritime risks described and conceptualised in ancient textual accounts? 
  • Can topic modelling and machine-learning classification detect systematic linguistic patterns associated with perceived risks, and, if so, do these patterns exhibit any non-causal correspondence with archaeological observations such as shipwreck densities? 
  • To what extent do textual representations of risk vary between regions characterised by distinct navigational hazards, and what might these variations reveal about local knowledge and seafaring strategies? 
  • How might integrating culturally embedded notions of risk into seafaring models refine our understanding of ancient navigational behaviour, particularly the assumption that the safest routes were necessarily the most probable ones? 

Methodology

Our study extends earlier work on seafaring strategies and risk perception in the Classical and Roman Mediterranean (Arnaud 2005) (Arnaud 2014) (Ritondale 2022) (Safadi / Sturt 2019) by incorporating natural language processing (NLP) based approaches to multilingual textual corpora associated with ancient seafaring. We combine direct (close) reading of key passages and navigational texts (e.g., periploi and portolans) with distant reading techniques capable of detecting lexical and semantic patterns across dispersed and heterogeneous datasets. This dual approach preserves the granularity of philological interpretation while enabling the identification of broader trends that remain inaccessible through traditional reading alone.

Currently, we are in the initial stages of data collation. Our contributions are as follows: (i) collate a dataset which contains texts describing a voyaging decision which mentions ports and anchorages within our areas of interest, and relevant information such as who wrote the text in what year; (ii) multilingual topic modelling on the collated dataset to observe topic trends relating to ancient seafaring over time; (iii) training a machine-learning classifier using textual embeddings to explore whether features in the text associated with perceived risks show any systematic relationship, without implying causation, with shipwreck incidences near the sites in question; and (iv) evaluating the trained classifier using explainable AI techniques to closely examine which textual aspects most strongly influence the classification, thus showing potential overlap between narrative constructions of risk, archaeological indicators, and actual environmental hazards.

Our corpus concentrates on three regions of strategic importance in Mediterranean seafaring, selected for their well-attested navigational hazards, their long-standing intensity of traffic, or a combination of both: the Gulf of Lion (France), the Italian Tyrrhenian coast (Excluding Sardinia) and the area from Cap Bon (Tunisia) to Alexandria (Egypt). For each region, we identify relevant passages from Greek and Latin texts in which a coastal site, anchorage or port is mentioned in the context of a voyaging decision. Further, we use the voyaging decision of sailing to a particular port or anchorage as a case study. For these sites, we extract texts in their original language as referenced in Pleiades (Bagnall et al. 2016). We will use a sentence-piece model trained on Latin and Ancient Greek (Riemenschneider / Frank 2023a) via knowledge distillation and BERTopic (Grootendorst 2022) to conduct topic modelling and extract interpretable topic clusters. For classification, we will use an encoder-only RoBerta-based model for Latin and Ancient Greek (Riemenschneider / Frank 2023b) and fine-tune it to classify the level of shipwrecks (low, medium or high) associated with the port mentioned in the texts. The trained classifier can then be explained using Shapley values (Lundberg et al. 2020) to extract key phrases or structures which contribute to the model accuracy. By following the outlined methodology, we aim to reconstruct attitudes towards maritime dangers while accounting for the cultural and temporal aspects of the corpora.

Part of the study is also devoted to reflecting on the methodological challenges inherent in applying NLP and AI techniques to ancient sources. Such approaches must contend with several issues, including textual preservation, semantic shift, translation bias, limited training data for historical languages, and toponymic ambiguities that introduce uncertainties and may affect the reliability of the outcomes. By explicitly engaging with these limitations and integrating computational analysis with domain-specific philological and archaeological expertise, the study does not seek to replace traditional approaches but to complement them. In doing so, it offers a critical framework for incorporating culturally embedded notions of risk into future seafaring models, thereby challenging prevailing assumptions and opening new avenues for understanding ancient navigational behaviour.

References
  1. Pascal, Arnaud (2005): Les routes de la navigation antique. Paris: Editions Errance.
  2. Pascal, Arnaud (2014): “Ancient mariners between experience and common sense geography.”, in: Features of common sense geography: Implicit knowledge structures in ancient geographical texts. Zürich: Lit Verlag 39–68.
  3. Bagnall, Roger / et al. (2016): “Pleiades: A gazetteer of past places.” <http://pleiades.stoa.org/> [06.05.2026].
  4. Grootendorst, Maarten (2022). “BERTopic: Neural topic modelling with a class-based TF-IDF procedure.” <arXiv preprint arXiv:2203.05794> [06.05.2026].
  5. Lundberg, Scott M. / Erion, Gabriel / Chen, Hugh / DeGrave, Alex / Prutkin, Jordan M. / Nair, Bala / Katz, Ronit / Himmelfarb, Jonathan / Bansal, Nisha / Lee, Su-In (2020): “From local explanations to global understanding with explainable AI for trees.”, in: Nature Machine Intelligence 2, 1:56-67.
  6. Maarleveld, Thijs J. (2004): “Finding ‘new’ boats: Enhancing our chances in heritage management, a predictive approach.”, in: Clark, Peter (eds.): The Dover Bronze Age boat in context: Society and water transport in prehistoric Europe. Oxford: Oxbow Books, 138-147.
  7. Riemenschneider, Frederick / Frank, Anette (2023a): “Graecia capta ferum victorem cepit. Detecting Latin Allusions to Ancient Greek Literature.”, in: Proceedings of the Ancient Language Processing Workshop, pages, Varna, Bulgaria: 30–38.
  8. Riemenschneider, Frederick / Frank, Anette (2023b): “Exploring Large Language Models for Classical Philology.”, in: Association for Computational Linguistics (ed.): Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics 1, Toronto, Canada: 15181–15199.
  9. Ritondale, Manuela (2022): Shipwrecking Probability in Mediterranean Territorial Waters: a Cultural Approach to Archaeological Predictive Modelling. Ph.D. thesis. University of Groningen, IMT School for Advanced Studies Lucca < https://doi.org/10.33612/diss.219256868> [06.05.2026].
  10. Safadi, Crystal / Sturt, Fraser (2019): “The warped sea of sailing: Maritime topographies of space and time for the Bronze Age eastern Mediterranean.”, in: Journal of Archaeological Science 103, 1–15.
  11. Slayton, Emma / Jarriel, Katherine / Montenegro, Álvaro / Safadi, Crystal / Smith, Karl / Zaia, Sara (2025): “Seafaring and modelling”, in: Journal of Maritime Archaeology 20, 601–628 <https://doi.org/10.1007/s11457-025-09455-5> [06.05.2026]