DH 2026

Daejeon, July 27–31

Poster

Lost Time: AI-Driven Dating and Typology of Post-1540 Hebrew Manuscripts as a Memory Tool

Gila Prebor
Bar Ilan University, Israel · gila.prebor@biu.ac.il

Introduction: The Inaccessible Archive

The dating and classification of undated manuscripts represent one of the most persistent challenges in the historiography of the Hebrew book. While traditional paleography relies on the manual expertise of a few scholars, and existing digital projects like SfarData have successfully mapped manuscripts up to 1540 based on established methodologies (Beit-Arié, 2022), a significant historical "black hole" remains. The period following 1540, characterized by the advent of print, increased scribal mobility, and the diversification of regional scripts, lacks a comprehensive typological framework. Consequently, thousands of manuscripts from the early modern period remain undated and contextualized in archives. This research seeks to fill this gap by providing an analytical framework for predicting migration patterns and locations (Prebor et al., 2020) through ontological analysis (Zhitomirsky-Geffet et al., 2020).

Methodology: Metadata as Cultural Evidence

Our methodology treats MARC 21 metadata from the National Library of Israel’s "KTIV" collection not just as cataloging records, but as a structured proxy for the physical manuscript. This framework is designed to address three key Research Questions:

  • AI Application: How can machine learning enhance the analysis of MARC metadata specific to post-1540 Hebrew manuscripts?
  • Typological Discovery: How can unsupervised learning identify new typological features and "scribal fingerprints" in manuscript metadata?
  • Dating Accuracy: How can AI algorithms incorporate historical and cultural nuances for accurate dating?

To answer these questions, we employ an unsupervised learning framework (see Figure 1):

  • Feature Extraction: We extract specific subfields (e.g., origin field 7XX, script type 546) to serve as digital proxies for scribal practices.
  • K-means Clustering: We utilize this algorithm to identify latent "spatiotemporal fingerprints" by analyzing correlations between scripts, locations, and dates.

Human-in-the-Loop Validation: To ensure reliability and mitigate AI "hallucinations," the generated clusters are manually reviewed by paleographical experts to confirm their historical validity. This hybrid approach ensures that the AI serves as a precision tool for historical reconstruction.

References
  1. Beit-Arié, M. (2022). Hebrew Codicology: Historical and Comparative Typology of Medieval Hebrew Codices based on the Documentation of the Extant Dated Manuscripts until 1540 from a Quantitative Approach. https://doi.org/10.25592/uhhfdm.9349
  2. Dhali, M. A. (2016). Artificial Intelligence in Historical Document Analysis: Pattern Recognition and Machine Learning Techniques in the Study of Ancient Manuscripts with a Focus on the Dead Sea Scrolls. University of Groningen. https://doi.org/10.33612/diss.869247881
  3. Prebor, G., Zhitomirsky-Geffet, M., & Miller, Y. (2020). A new analytic framework for prediction of migration patterns and locations of historical manuscripts based on their script types. Digital Scholarship in the Humanities, 35(2), 441-458.
  4. Suissa, O., Elmalech, A., & Zhitomirsky-Geffet, M. (2022). Improving OCR errors correction in Hebrew historical documents using deep learning methods. Journal of Documentation, 78(7), 46-65.
  5. Wahlberg, F., Wilkinson, T., & Brun, A. (2016). Historical manuscript production date estimation using deep convolutional neural networks. In 2016 15th International Conference on Frontiers in Handwriting Recognition (ICFHR) (pp. 205-210). IEEE.
  6. Zhitomirsky-Geffet, M., Prebor, G., & Miller, I. (2020). Ontology-based analysis of the large collection of historical Hebrew manuscripts. Digital Scholarship in the Humanities, 35(3), 688-719.