DH 2026

Daejeon, July 27–31

Wed, July 2909:00–10:30S108Grand Ballroom
Short Paper

When Classics Meets AI: Participatory Research for Building an NLP Infrastructure

Andrea Beyer
Humboldt-Universität zu Berlin, Germany · beyeranz@hu-berlin.de

Introduction

Classical Philology or Classics is a discipline that works with predominantly literary texts in the historical languages Latin and Ancient Greek. In Germany, the majority of researchers focus on ancient sources usually preferring research areas such as editions and literary studies which use close reading (Greenham 2019). Meanwhile, it seems that they consider linguistics and teaching to be less important. In contrast, data-driven and AI research methods are more widely applied in Classics research globally (Vatri / McGillivray 2020; Sommerschield et al. 2023), even though they seem to have little effect on literary studies here (Hagel 2022; La Veglia 2024). This is also due to the fact that this research is generally not aimed at gaining new insights into literary studies, but rather at improving a method (McGillivray 2013; Assael et al. 2022), providing a data set (Dexter et al. 2024; Stopponi et al. 2024) or developing new algorithms (Sprugnoli et al. 2019; Burns 2023). Although the interest especially in the methods of Natural Language Processing (NLP) and the willingness to engage with them has grown overall (Shang et al. 2025), the disruption caused by Large Language Models (LLM) has yet not led to a noticeable increase in digitally supported research projects across the Classic’s community (La Veglia 2024). One main reason seems to be a significant gap between the individual research competence and the criteria requirements (Filograsso et al. 2025) – including skills in statistical and linguistic methods, data management and visualisation – associated with what is known as distant reading (Moretti 2013). So, there are two gaps: Firstly, NLP researchers do not consider the needs of classicists conducting research on literary texts. Secondly, classical philologists who favour qualitative methods lack the digital literacy (Buckingham 2010), data literacy (Schmidt et al. 2021) and AI literacy (Laupichler et al. 2022) necessary to engage with data-driven methods. To close both gaps, we have developed Daidalos, the web-based prototype of an NLP research infrastructure within a third-party funded interdisciplinary project (Beyer / Schulz 2024; Schulz / Kotschka, 2026). This prototype provides access to NLP models for literary studies (such as Word2Vec (Mikolov et al. 2013) for word embeddings or LatinAffectus (Sprugnoli et al. 2020) for sentiment analysis), as well as explanations and learning materials, via no-code (NLP tools with a graphical user interface), low-code (Jupyter notebooks) and code (application programming interface) options.

Methodological Approach

The idea of developing such an infrastructure originated in our engagement in and collaboration with a Community of Practice (CoP) of German speaking researchers, teachers, and students of Classical Philology. Based on the concepts of Situated Learning (Lave / Wenger 1991) and Participatory Research (PR) (Vaughn / Jacquez 2020; Ma et al. 2025) we have established research tandems which bring together experts and laypersons related to the methods of the Digital Humanities (DH): On the one hand, we use systematic inquiry in direct collaboration to develop and evaluate the research infrastructure, e. g. for user stories (Wautelet et al. 2017) and participatory design (Assis et al. 2025). On the other hand, we empower our research partners who generally lack training in DH methods but are deeply affected by the methodological shift in the humanities (Amangazykyzy et al. 2025), to embrace a new mix of research methods. If they then work as multipliers, as is intended, we can achieve the necessary long-term methodological change in the Classics together.

Findings

After two years of project work and applying the PR approach we can provide the following insights:

  • Technical infrastructure: Agile software development and PR are highly synergistic. The prototype was successfully built based on user stories and was repeatedly tested by different members of the CoP who required a more intuitive guidance through the research workflow, among other things.
  • Social infrastructure: PR and a social infrastructure offering networks, workshops, consulting, and open educational resources are inherently connected. This approach is highly appreciated by the CoP and helps its members to become more confident and equal in their understanding of NLP. It thrives on interaction and must be specifically targeted at individuals or interest groups to encourage their engagement and willingness to change their established research behaviours.
  • Software developers: PR is extremely helpful in understanding the settings and ways in which NLP methods are important in everyday research on Latin and Ancient Greek texts. Nevertheless, working across disciplines and their subordinate research areas is challenging and is best countered by an interdisciplinary developer group that brings together different vocabularies and various methods while sharing important values, e. g. the equality of diverse partners.
  • CoP: Although PR is not common in the addressed CoP, it was welcomed. It helped us identify major challenges in opening up the CoP to DH methods. For example, we overestimated the existing digital research competence, e. g. interpreting visualisations of text coverage for sentiment lexicon or of association network graphs for a trained word embeddings model. Therefore, changes in research methodology cannot be addressed without also emphasising teaching and learning.
  • Impact: PR and long-term outcomes, such as changes to a CoP, take time. Thus, a project restricted to a few years does not have enough leverage to change long-established habits. While it can conceptualise and initiate a transformation, long-term funding is needed for the implementation and evaluation of a change.

Conclusion

PR is well suited to building a research infrastructure that aims to provide the means to change research practices. However, it is a more suitable approach for an institutionalised context because it is time-consuming and requires a longer timeframe than a project allows for. Moreover, if applied to a research community, PR raises some critical issues, e. g. the academic imbalance between the developers and the sometimes very renowned researchers or the lack of funding to improve curricula and teaching. In general, this approach is not scalable enough to facilitate widespread change in a CoP unless it is adopted by the community itself, demonstrating a strong commitment to lifelong learning and allocating resources to support human-centred infrastructures for research and learning.

References
  1. Amangazykyzy, Moldir / Gilea, Aigerim / Karlygash, Aubakirova / Nurziya, Abisheva / Sandygash, Kulanova (2025): “Epistemological Transformation of the Paradigm of Literary Studies in the Context of the Integration of Digital Humanities Methods”, in: Forum for Linguistic Studies 7, 4: 166–176. DOI: 10.30564/fls.v7i4.8619.
  2. Assael, Yannis / Sommerschield, Thea / Shillingford, Brendan / Bordbar, Mahyar / Pavlopoulos, John / Chatzipanagiotou, Marita / Androutsopoulos, Ion / Prag, Jonathan / de Freitas, Nando (2022): “Restoring and attributing ancient texts using deep neural networks”, in: Nature 603, 7900: 280–283. DOI: 10.1038/s41586-022-04448-z.
  3. Assis, João Victor / Guerino, Guilherme / Rodrigues, Luiz / Macario, Valmir / Marinho, Marcelo (2025): “Integrating Participatory Design and Dual-track Software Development Process: A Case Study From an Intelligent Mathematics Tutoring System”, in: Congresso Ibero-Americano em Engenharia de Software (CIbSE), May 2025: 30–44. DOI: 10.5753/cibse.2025.35290.
  4. Beyer, Andrea / Schulz, Konstantin (2024): “Daidalos-Projekt - Entwicklung einer Infrastruktur zum Einsatz von Natural Language Processing für Forschende der Klassischen Philologie”, in: Zenodo. DOI: 10.5281/zenodo.12635794.
  5. Buckingham, David (2010): “Defining Digital Literacy”, in: Bachmair, Ben (ed.): Medienbildung in neuen Kulturräumen. Wiesbaden: Verl. für Sozialwissenschaften 59–71. DOI: 10.1007/978-3-531-92133-4‗.
  6. Burns, Patrick J. (2023): “LatinCy: Synthetic Trained Pipelines for Latin NLP”, in: arXiv https://arxiv.org/pdf/2305.04365.pdf [03.05.2026].
  7. Dexter, Joseph P. / Chaudhuri, Pramit / Burns, Patrick J. / Adams, Elizabeth D. / Bolt, Thomas J. / Cásarez, Adriana / Flynt, Jeffrey H. / Li, Kyle / Patterson, James F. / Schwartz, Ariane / Shumway, Scott (2024): “A Database of Intertexts in Valerius Flaccus’ Argonautica: A Benchmarking Resource for the Evaluation of Computational Intertextual Search of Latin Corpora”, in: Journal of Open Humanities Data 10, 1. DOI: 10.5334/johd.153.
  8. Filograsso, Francesca / Massari, Arcangelo / Neri, Camillo / Peroni, Silvio (2025): “HERITRACE in action: the ParaText project as a case study for semantic data management in Classical Philology”, in: arXiv http://arxiv.org/abs/2508.15556 [02.12.2025].
  9. Greenham, David (2019): Close reading: The basics. First edition. Boca Raton, FL: Routledge.
  10. Hagel, Stefan (2022): “How Is Technology Useful in the Study of Ancient Music?”, in: Greek and Roman Musical Studies 10, 2: 269–289. DOI: 10.1163/22129758-bja10043.
  11. La Veglia, Andrea (2024): “Being a Classicist in the Digital Age”, in: Reggiani, Nicola (ed.): Digital Papyrology III: The Digital Critical Edition of Greek Papyri: Issues, Projects, and Perspectives. Berlin, Boston: De Gruyter 49–70. DOI: 10.1515/9783111070162.
  12. Laupichler, Matthias Carl / Aster, Alexandra / Schirch, Jana / Raupach, Tobias (2022): “Artificial intelligence literacy in higher and adult education: A scoping literature review”, in: Computers and Education: Artificial Intelligence 3: 100101. DOI: 10.1016/j.caeai.2022.100101.
  13. Lave, Jean / Wenger, Etienne (1991): Situated Learning: Legitimate Peripheral Participation. Cambridge: Cambridge University Press. DOI: 10.1017/CBO9780511815355.
  14. Ma, Rongqian / Chen, Annie T. / Bossaller, Jenny / Boyles, Christina / Donaldson, Devan Ray (2025): “Co‐Creation in Context: Participatory Approaches to Digital Humanities and Cultural Heritage Work”, in: Proceedings of the Association for Information Science and Technology, Washington D.C., 2025: 1264–1269. DOI: 10.1002/pra2.1379.
  15. McGillivray, Barbara (2013): Methods in Latin computational linguistics. Leiden / Boston: Brill.
  16. Mikolov, Tomas / Chen, Kai / Corrado, Greg / Dean, Jeffrey (2013): “Efficient Estimation of Word Representations in Vector Space”, in: arXiv preprint arXiv:1301.3781.
  17. Moretti, Franco (2013): Distant Reading. Verso Books.
  18. Schmidt, Andreas / Neifer, Thomas / Haag, Benedikt (2021): “Data Literacy als ein essenzieller Skill für das 21. Jahrhundert”, in: Frick, Detlev / Gadatsch, Andreas / Kaufmann, Jens / Lankes, Birgit / Quix, Christoph / Schmidt, Andreas / Schmitz, Uwe (eds.): Data Science: Konzepte, Erfahrungen, Fallstudien und Praxis. Wiesbaden: Springer Fachmedien 27–40. DOI: 10.1007/978-3-658-33403-1_2.
  19. Schulz, Konstantin / Kotschka, Florian (2026): Daidalos. Berlin https://daidalos-projekt.de/ [03.05.2026].
  20. Shang, Wenyi / Ma, Rongqian / Moulaison-Sandy, Heather (2025): “How does digital humanities research talk about AI? A bibliometric analysis”, in: Information Research. An international electronic journal 30, iConf: 635–645. DOI: 10.47989/ir30iConf47242.
  21. Sommerschield, Thea / Assael, Yannis / Pavlopoulos, John / Stefanak, Vanessa / Senior, Andrew / Dyer, Chris / Bodel, John / Prag, Jonathan / Androutsopoulos, Ion / de Freitas, Nando (2023): “Machine learning for ancient languages: A survey”, in: Computational Linguistics 49, 3: 703–747. DOI: 10.1162/coli_a_00481.
  22. Sprugnoli, Rachele / Passarotti, Marco / Corbetta, Daniela / Peverelli, Andrea (2020): “Odi et Amo. Creating, Evaluating and Extending Sentiment Lexicons for Latin.”, in: Proceedings of The 12th Language Resources and Evaluation Conference: 3078–3086 https://aclanthology.org/2020.lrec-1.376.pdf [03.05.2026].
  23. Sprugnoli, Rachele / Passarotti, Marco / Moretti, Giovanni (2019): “Vir is to Moderatus as Mulier is to Intemperans. Lemma Embeddings for Latin”, in: CLiC-it https://ceur-ws.org/Vol-2481/paper69.pdf [03.05.2026].
  24. Stopponi, Silvia / Peels-Matthey, Saskia / Nissim, Malvina (2024): “AGREE: a new benchmark for the evaluation of distributional semantic models of ancient Greek”, in: Digital Scholarship in the Humanities. DOI: 10.1093/llc/fqad087.
  25. Vatri, Alessandro / McGillivray, Barbara (2020): “Lemmatization for Ancient Greek: An experimental assessment of the state of the art”, in: Journal of Greek Linguistics 20, 2: 179–196. DOI: 10.1163/15699846-02002001.
  26. Vaughn, Lisa M. / Jacquez, Farrah (2020): “Participatory Research Methods – Choice Points in the Research Process”, in: Journal of Participatory Research Methods 1, 1. DOI: 10.35844/001c.13244.
  27. Wautelet, Yves / Heng, Samedi / Kiv, Soreangsey / Kolp, Manuel (2017): “User-story driven development of multi-agent systems: A process fragment for agile methods”, in: Computer Languages, Systems & Structures 50: 159–176. DOI: 10.1016/j.cl.2017.06.007.