Daejeon, July 27–31
This short paper introduces the CLARUS Multilingual Lexicon of Digital Forensics, a searchable web-based resource developed through a systematic corpus-driven methodology to address terminological ambiguity in cross-border miscommunication in digital forensic practice. The lexicon currently comprises 212 terms/entries across five languages (English, Czech, Finnish, Greek, and Portuguese) which are drawn from patterns in actual professional usage. Unlike existing resources such as SWGDE's terminology standards and NIST's OSAC Lexicon (which adopt prescriptive approaches defining how terms should be used), our corpus-driven methodology captures how practitioners actually employ terminology in professional contexts. This project aligns with the DH2026 conference theme of 'Engagement' by fostering meaningful connections between diverse linguistic communities and addressing technological challenges in international forensic collaboration.
Digital forensic investigations increasingly transcend national boundaries, requiring collaboration between practitioners operating within different legal frameworks, organisational cultures, and linguistic contexts. Recent systematic reviews have demonstrated miscommunication and terminological confusion may affect forensic decision-making across multiple disciplines (Cooper & Meterko 2019; Kukucka et al. 2017).
While existing forensic glossaries have been developed by organisations such as the Scientific Working Group on Digital Evidence (SWGDE) and the National Institute of Standards and Technology's Organization of Scientific Area Committees (NIST/OSAC), these resources predominantly adopt prescriptive approaches that standardise how terms should be used by practitioners. However, practitioners often employ terminology differently in actual practice, creating potential for misunderstandings in cross-border contexts where multiple jurisdictional standards intersect.
This gap between prescriptive terminology and actual usage patterns presents both a linguistic challenge and an opportunity for the digital humanities. Our response was to develop a corpus-driven multilingual lexicon grounded in empirical evidence of how digital forensic practitioners actually use language, applying established corpus linguistic methodologies (McEnery & Hardie 2012; Baker 2010) to specialised professional discourse.
The lexicon development employed a corpus linguistic methodology adapted for forensic practice, involving a five-step methodological framework (see Candarli & Balfour 2026 for details): (1) building a specialised corpus of over 1,000,000 words in the domain of digital forensics; (2) extracting key terms for possible inclusion in the lexicon; (3) refining candidate terms through a questionnaire administered to practitioners; (4) writing definitions by examining concordance lines; and (5) reviewing and refining the definitions. This methodology is fully replicable and transferable to other forensic evidence domains or specialised professional fields requiring terminological standardisation, aligning with ISO principles for multilingual terminology management (ISO 12616-1:2021).
Each of the 212 entries includes: multilingual equivalents in five languages, professionally translated to ensure accuracy of terminology; semantic categorisation (digital evidence, forensic methods, digital technology, forensic infrastructure, legal frameworks) reflecting domains in which the terms might be used; corpus-derived definitions grounded in actual professional usage; collocation patterns showing typical word combinations; frequency indicators (low/medium/high) gauging prevalence of the term in forensic literature; and usage examples in all languages.
The web-based interface enables practitioners to search by term, semantic category, or language. The lexicon and framework documentation will be freely available online, allowing the lexicon to have global reach and encourage further refinement.
The lexicon functions as a critical digital resource for supporting international forensic cooperation. By providing validated multilingual equivalents and usage patterns, it reduces the risk of miscommunication that could contribute to the cognitive biases and errors documented in forensic science literature. The documented five-step methodology enables expansion of the lexicon to other evidence types (biological, financial, physical evidence) and additional languages, creating potential for additional lexicons in digital forensics (see Balfour & Candarli 2026 for the framework of expanding the lexicon).