Daejeon, July 27–31
Digital work in Turcology has expanded through corpus resources, digital lexicography, dialect atlases, digital editions, network analysis, digital history, manuscript digitisation, NLP resources, and AI-oriented experiments. Corpus and NLP-oriented developments are represented by the Turkish National Corpus and Turkish NLP resource surveys (Aksan et al. 2012; Çöltekin et al. 2023), while digital lexicography, collocation, and term-extraction work has contributed to this landscape (Tahiroğlu 2010). Examples include: stylometric authorship attribution has examined disputed literary authorship in the Reşat Enis case (Çetin Baycanlar / Tahiroğlu 2022), and AI-generated visual analysis has been applied to colour–emotion relations in Stable Diffusion illustrations of animal figures from Turkish mythology (Yıldırım / Yılmaz Arıkan 2026). These practices are valuable, but dispersed, and have been described through labels such as “digital Turkish”, “digital Turcology”, “computer-assisted Turcology”, and “computational Turcology”. This creates the need for shared terminology, evaluation criteria, and a clearer account of how linguistic, textual, historical, material, and cultural evidence can be modelled, documented, reused, and governed (Borgman 2015; Wilkinson et al. 2016).
This short paper proposes Digital Turcology and Cultural Heritage Studies (DTCHS) as an emerging field of specialisation and offers an engagement-centred framework for connecting dispersed digital practices across Turcology. Turcology is used in its interdisciplinary sense, encompassing language, literature, history, history of science, art history, folklore, cultural practices, and heritage formations across periods, scripts, and regions. In the age of Digital Humanities and computational methods, these areas face shared questions: how to structure evidence, model interpretation, document uncertainty, design reusable data, build tools, and make cultural knowledge computable while preserving historical complexity.
DTCHS responds as a DH-grounded and computationally informed field. It brings Digital Humanities methods, cultural heritage data practices, and NLP, machine learning, and AI-assisted workflows into dialogue with the evidentiary needs of Turkish and Turkic materials. Its central methodological proposition is engagement as a method. Engagement is defined here as the structured involvement of domain experts, communities, users, students, data stewards, technical collaborators, and future reusers in the design, documentation, validation, evaluation, and reuse of digital research outputs. This formulation draws on participatory approaches to cultural institutions while extending engagement into digital research infrastructures (Simon 2010). As a research method and criterion of quality, engagement shapes research questions, evidence encoding, uncertainty documentation, terminology governance, tool and workflow design, data modelling, and long-term reuse. It also requires scholars, collections, languages, and cultural communities from under-represented Turkic contexts to be treated as co-designers of research questions, data models, tools, and criteria.
The need for such a framework becomes visible at points of friction between general-purpose DH tools, NLP and machine-learning pipelines, AI systems, network-analysis workflows, and Turkish or Turkic linguistic realities. These include Arabic-script Turkish materials, historical transcription and transliteration, agglutinative morphology, multilingual technical vocabularies, low-resource Turkic languages, manuscript-based evidence, oral and material heritage, and online cultural traces whose meaning depends on platform-specific circulation. They intersect with wider problems of linguistic diversity and uneven resource distribution in NLP (Joshi et al. 2020), while Turkish NLP resource surveys show how tool coverage and processing assumptions shape what can be computationally handled in Turkish data (Çöltekin et al. 2023). In the author’s TEI-based work on Ottoman mathematical manuscript material, such problems become visible in transcription, transliteration, mathematical notation, terminology, and editorial documentation (Kalafat 2024). A further issue is the inconsistent handling of Turkish-specific characters such as ç, ğ, ı, ö, ş, and ü in data cleaning, tokenisation, graph labels, export formats, and AI-assisted workflows. Forced ASCII normalisation can turn meaningful lexical forms into distorted tokens. For example, the historical Turkish mathematical term sagış/sağış, meaning “calculation” or “number”, may be reduced to sagis, obscuring orthographic, phonological, and terminological continuity. TEI, IIIF, PROV, FAIR, CIDOC CRM, and terminology and ontology practices provide adaptable infrastructures for representing evidence, relationships, and provenance (Doerr 2003; W3C 2013a; W3C 2013b; Wilkinson et al. 2016; IIIF Consortium 2020; TEI Consortium 2025).
The framework is evaluated through six dimensions: interpretability of evidence, traceability of analytical and editorial decisions, terminology governance, documentation of uncertainty and plurality, interoperability of data models, and sustainability beyond the original project. Evaluation concerns whether an evidence chain can be inspected, reused, extended, questioned, and taught. These dimensions are informed by, but not reducible to, provenance, FAIR-oriented reuse, and TEI-based evidence representation (W3C 2013a; W3C 2013b; Wilkinson et al. 2016; TEI Consortium 2025).
Against this evaluative framework, the paper demonstrates the approach through three testbeds drawn from the author’s work. The first concerns TEI-based digital historical text editing for Turkish manuscript materials, grounded in work on the Ottoman mathematical manuscript Risâle-i Sinüs (Kalafat 2024). Building on this experience, the ongoing Türkî Hisâb work develops a TEI-first approach to historical terminology and digital philological analysis, focusing on context–term–grammar relations and a TEI-linked pilot termbank for specialised Turkish mathematical vocabulary. The second testbed is a digital cultural semiotic network of Turkish wool-knitting motif names, based on a repertoire of approximately 750 motif names and a subcorpus of 16 motifs naming women’s affinal kinship roles (Kalafat 2026a). The third is a cultural heritage data science micro-corpus tracing the online circulation and memetic transformation of a Turkish idiom (Kalafat 2026b).
Within this framing, Ottoman materials form a significant domain within Turcology. Much recent digital work on Ottoman materials has developed around Digital Ottoman Studies (Barakat / Yaycıoğlu 2022; Aladağ 2024). Here, this line of work is read primarily as history-oriented, archive-oriented, or tool-oriented digital scholarship. DTCHS places philological evidence chains at the centre: transcription, transliteration, terminology, textual variation, provenance, and the long-term history of Turkish intellectual vocabulary. Digital Ottoman Studies is therefore positioned as a tool-oriented subdomain within DTCHS, rather than as a founding layer or separate umbrella field.
The paper introduces the digital turcologist as a practitioner profile combining Turcological expertise with data literacy, computational thinking, DH-tech practice, workflow design, tool-building, knowledge modelling, and evaluation. Conceptually, DTCHS is an emerging field broad enough to encompass Turcology, yet specific enough to guide method. Methodologically, it connects DH, digital philology, cultural heritage data science, computational, NLP/ML, and AI-assisted research through evidence, interoperability, terminology governance, and accountability.
Keywords
DTCHS, engagement, interoperability, translinguality, cultural heritage data