Daejeon, July 27–31
Introduction: The Inaccessible Archive
The dating and classification of undated manuscripts represent one of the most persistent challenges in the historiography of the Hebrew book. While traditional paleography relies on the manual expertise of a few scholars, and existing digital projects like SfarData have successfully mapped manuscripts up to 1540 based on established methodologies (Beit-Arié, 2022), a significant historical "black hole" remains. The period following 1540, characterized by the advent of print, increased scribal mobility, and the diversification of regional scripts, lacks a comprehensive typological framework. Consequently, thousands of manuscripts from the early modern period remain undated and contextualized in archives. This research seeks to fill this gap by providing an analytical framework for predicting migration patterns and locations (Prebor et al., 2020) through ontological analysis (Zhitomirsky-Geffet et al., 2020).
Methodology: Metadata as Cultural Evidence
Our methodology treats MARC 21 metadata from the National Library of Israel’s "KTIV" collection not just as cataloging records, but as a structured proxy for the physical manuscript. This framework is designed to address three key Research Questions:
To answer these questions, we employ an unsupervised learning framework (see Figure 1):
Human-in-the-Loop Validation: To ensure reliability and mitigate AI "hallucinations," the generated clusters are manually reviewed by paleographical experts to confirm their historical validity. This hybrid approach ensures that the AI serves as a precision tool for historical reconstruction.