DH 2026

Daejeon, July 27–31

Thu, July 3013:40–15:10S018105
Long Paper

Reconstructing Legal Knowledge Across Time: \\SKOS--OWL Modeling, Semantic Drift, and AI-Augmented Translation \\of Mkhitar Gosh’s Medieval Law Code

Gayane Hovhannisyan
European University of Armenia, Armenia · g.hovhannisyan@eua.am
Hamest Tamrazyan
EPFL/Switzerland, Switzerland · hamest.tamrazyan@epfl.ch

Introduction

Medieval legal texts seldom appear in discussions of semantic modeling or multilingual AI, yet they represent some of the most resilient epistemic structures in human intellectual history (Byrne 2020). Their endurance rests on a productive tension: a conceptual core that persists over centuries, and interpretive edges that evolve through linguistic change, localized commentary, and new social imaginaries. This paper examines Mkhitar Gosh’s Law Code (1184) in its 1880 and 1975 versions, as well as a UNESCO Memory of the World manuscript (UNESCO 2025). Over eight centuries, the Law Code circulated in manuscript form. It was glossed by scribes, translated into Armenian vernaculars and major Eurasian languages, and adapted to evolving legal and theological contexts (Baronch 1869; Wucicki 1843; Kod 1828; Avagyan 2001; Zbi 1906). This long transmission history provides a unique opportunity to observe how legal meaning travels across time and translation and to reconstruct its conceptual structure through SKOS vocabularies, OWL ontologies, and AI-assisted translation analysis. We argue that modeling medieval law is not a matter of imposing modern computational logic onto historical text. Instead, it requires recovering the epistemological architecture of Classical Armenian legal reasoning first, and then translating it into interoperable digital structures that preserve cultural specificity. A hybrid semantic approach, particularly SKOS for conceptual precision and OWL for legal logic, allows us to capture the interpretive flexibility and structural coherence of Gosh’s legal system. This aims to integrate minority legal traditions into multilingual DH infrastructures and machine translation workflows.

Research Questions

Our study addresses four guiding questions:

Representation: How can the conceptual categories of a medieval legal text be modeled without flattening their cultural values?

Semantic Drift: What shifts emerge when Classical Armenian clauses are compared with modern Armenian renderings, human English translations, and AI- generated versions?

Ontology + AI: How can SKOS and OWL jointly support culturally respectful, interpretable AI-assisted translation for low-resource legal languages?

Design: What does the Law Code’s transmission history contribute to building inclusive, interoperable legal infrastructures for underrepresented languages?

Methodology and Corpus

Corpus and Data

Our corpus includes:

  • the Classical Armenian (Grabar) text of Gosh’s Law Code (Gosh 1880);
  • modern Armenian editions (Gosh 1975);
  • the authoritative English translation (Thomson 2000);
  • AI-generated modern English translations (GPT-based);
  • manually curated SKOS and OWL files (Hovhannisyan 2025).

The combination of philological, semantic, and computational layers enables a multi- lens analysis of how legal meaning is structured, transmitted, and reinterpreted.

Methodology

The methodology integrates historical semantics, ontology engineering, translation analysis, and AI-assisted modeling within a structured workflow. Eight clauses were selected from the eight legal domains covered by the Code. AI translations were generated using ChatGPT Pro (GPT-5.2, Summer 2025) via the web interface. Decoding parameters were set to ensure terminological stability (temperature = 0.2; top-p = 0.9). The following workflow comprises a. clause segmentation and selection; b. manual extraction of operative legal concepts; c. SKOS concept modeling and hierarchical structuring; d. OWL formalization of logical relations; e. cross-translation alignment. Final outputs were reviewed and refined through ontology-guided post-editing.

Conceptual Extraction

We identify the conceptual primitives embedded in the Law Code:

  • Legal actors: judge, accuser, heir
  • States: intent, culpability
  • Actions: oath-breaking, theft
  • Sanctions: restitution, compensation
  • Principles: proportionality, authority (Hovhannisyan 2025).

These categories organize the text's logic. For example, the Classical Armenian clause: « Վարձուք ըստ գործոցն իւրեանց սահմանեցան » (“Compensation is established according to one’s deeds”) expresses proportionality as a moral–legal axiom, not a punitive doctrine. Many terms, such as մեղք (sin), span theological, moral, and legal contexts. Conceptual detail is, therefore, essential to the fidelity of digital models; we treat lexical items as cultural concepts rather than as mere translation targets.

SKOS Conceptual Modeling

SKOS preserves conceptual nuance by supporting hierarchical relations, associative links, multilingual labels, and culturally grounded definitions (W3C 2009; Baker et al. 2013). For example, the term ապաշխարություն (“penance”) spans legal, moral, and theological domains; SKOS preserves these layers without collapsing them into a single category. This flexibility is essential for representing the multivalent nature of medieval legal concepts.

OWL Legal Logic Modeling

OWL was used to express the Law Code’s conditional and hierarchical reasoning (Ceci and Gangemi 2016; Schneider and Sutcliffe 2011). In Article 83 of Gosh (1880) we read: “If the pledge is lost through the negligence of the pledge-holder, he shall repay it fourfold.”. We model this clause with:

  • Class: PledgeLoss
  • Properties: hasCause some Negligence; hasAgent some PledgeHolder;
  • Consequence: hasPenalty some FourfoldCompensation.

This transforms a medieval clause into a reasoning-ready structure—a computational analog of its logic. We use OWL-DL for computability and OWL-Full or punning when SKOS concepts need to function as both classes and individuals.

Comparative Translation Analysis

To trace semantic drift, we aligned Classical Armenian → modern Armenian → human English → AI-generated English translations. The clause « Վարձուք ըստ գործոցն իւրեանց սահմանեցան »(literally: “Compensation is determined according to the deeds.”) becomes “Let the punishment be according to the deeds” (Thomson 2000) and in AI output “Punishment should be proportionate to the offense.” The shift from restitution to punitive proportionality, and from “deeds” to “offense,” illustrates how conceptual nuance erodes across translation layers.

Similarly, the idiom լինել ի վերայ անդատաստանայ (“to stand above judgment”) shifts from a moral criterion in Classical Armenian to procedural authority in modern Armenian and to an administrative function in English and AI translations. These divergences indicate where ontological grounding is required.

Human-AI Complementarity

We adopt the principle of complementarity: human translators preserve nuance; AI identifies structural regularities; ontologies stabilize meaning across both. AI tends to modernize legal lexicon and simplify moral logic, reinforcing the need for conceptual scaffolding. Our semantic model serves as the shared reference frame that ties probabilistic outputs to historical reasoning.

Results and Discussion

Semantic Modelability of the Law Code

The Law Code’s structure, case-by-case scenarios with explicit agents, conditions, actions, and outcomes, mirrors the logic of semantic web modeling, including conceptual containers of roles, context or conditions, actions and consequences.

Conceptual Stability Across Time and Translation

Despite linguistic variation over the centuries, the Law Code’s conceptual core appears largely stable. Distinctions between intentional and unintentional harm persist; restitution consistently remains the preferred legal response; judicial authority retains its blended clerical–administrative character; and the ethical significance of “deeds” never disappears. Each clause functions as a micro-ontology, making explicit the relationships and constraints that can be encoded in OWL.

Translation Drift Reveals Hidden Legal Structures

Aligning translations reveals where the Law Code’s implicit logic becomes explicit in modern renderings. The idiom “to stand above judgment” becomes “to preside over judgment” in human English and “to render justice and pronounce rulings” in AI output. These shifts show how moral concepts transform when mapped onto modern legal lexicons. AI exaggerates this drift by normalizing medieval phrasing into contemporary legal English. Such differences indicate which specific linguistic and logical categories must be intentionally maintained within the digital knowledge system.

SKOS and OWL Serve Complementary Functions

Our modeling shows that, on the one hand, SKOS handles culturally embedded, fuzzy, hierarchical relations. On the other hand, OWL handles strict legal logic and conditional constraints. Neither is sufficient on its own for historical legal material. Together, they capture both worldview and rule structure. For instance, the category “boundary dispute” is both a concept in SKOS (linked to property, neighbor, oath) and a class in OWL (requiring propertyOwner, boundaryMarker, and disputeEvent) (W3C 2009; Baker et al. 2013; Ceci and Gangemi 2016; Schneider and Sutcliffe 2011).

Both layers are needed to systematically capture historical legal reasoning: OWL enforces structure, and SKOS preserves cultural nuance.

Ontology-Guided MT for Low-Resource Languages

After the ontology-guided post-editing of the translation, AI translation becomes more terminologically consistent, less prone to anachronism, avoids semantic flattening of cultural detail, is more sensitive to hierarchical role distinctions, and handles theological–legal hybrids more reliably. This is essential not only for Armenian but also for languages with limited training data in broader applications.

Semantic Modeling as Digital Manuscript Modeling

Just as medieval scribes clarified concepts, reorganized clauses, added glosses, and adapted content to new contexts, our semantic model performs this mediating function for the digital era. Modeling becomes a continuation of the Law Code’s historical life.

Contribution to Digital Humanities

This research contributes to DH by offering a methodology for culturally grounded semantic modeling of legal heritage. Second, it demonstrates how SKOS/OWL integration can represent complex multilingual legal corpora. Moreover, it provides a transferable workflow for low-resource languages, enriching AI-assisted translation with domain-specific conceptual grounding, promoting FAIR, interoperable legal ontologies, and showing how DH methods can preserve both semantic precision and cultural worldview. Crucially, our work positions legal heritage as computationally meaningful without reducing its historical complexity.

Conclusion

The proposed semantic model of Gosh’s Law Code demonstrates that medieval legal texts can be systematically represented within contemporary computational frameworks while preserving their cultural specificity. By integrating SKOS, OWL, RDF, philological analysis, and AI-assisted translation, the study reconstructs key elements of the legal epistemology encoded in Classical Armenian and renders them interoperable with Digital Humanities infrastructures. More broadly, the case indicates that cultural heritage, often resistant to computational formalization, can contribute culturally embedded structured logic to AI-mediated knowledge systems. At the same time, the findings caution against assumptions of universal legal equivalence. Although full cross-cultural equivalence cannot be presumed in legal translation, ontology-based SKOS/OWL mediation improves semantic transparency and facilitates the integration of historical, underrepresented, and endangered legal corpora into evolving digital knowledge environments.

References
  1. Kodeks zakonov Armyanskikh. (1828). Tipografiya Karla Kraya, St Petersburg.
  2. Zbiór Prawa Polskiego, volume 3. (1906). Krakowska Składnica Pomocy Naukowych, Kraków. Avagyan, R. (2001). Hye iravakan mtqi gandzaran [The Treasury of Armenian Legal Thought], Volume 2. MNUI-XXI, Yerevan.
  3. Baker, T., Bechhofer, S., Isaac, A., Miles, A., Schreiber, G., and Summers, E. (2013). Key choices in the design of a simple knowledge organization system (SKOS).
  4. Baronch, A. (1869). Kodeks praw Ormian polskich Mychitara Gosza. Drukarnia Zakładu Naro- dowego im. Ossolińskich, Lwów.
  5. Byrne, P. (2020). Medieval violence, the making of law and the historical present. Journal of the British Academy.
  6. Ceci, M. and Gangemi, A. (2016). An owl ontology library representing judicial interpretations (judo). Semantic Web, 7(3):229-253.
  7. Gosh, M. (1880). Girq Datastani [The Law Code]. Arsen Bastamyants, Venice.
  8. Gosh, M. (1975). Girq Datastani [The Law Code]. Armenian SSR Academy of Sciences, Yerevan.
  9. Hovhannisyan, G. (2025). Dataset of Mxitar Gosh's law code (Armenian and English).
  10. Memory of the World Register (UNESCO 2025). Law Code of Mkhitar Gosh.
  11. Schneider, M. and Sutcliffe, G. (2011). Reasoning in the OWL 2 full ontology language using first- order automated theorem proving.
  12. Thomson, R. W. (2000). The Law Code (Datastanagirkʿ) of Mxitʿar Goš. Rodopi, Amsterdam.
  13. W3C (2009). SKOS simple knowledge organization system: core vocabulary specification. Technical report.
  14. Wucicki, K. (1843). Zbiór praw ormiańskich z łamanego języka ormiańskiego na polski przełożony. Drukarnia Stanisława Strąbskiego, Warsaw.