DH 2026

Daejeon, July 27–31

Fri, July 3109:00–10:30S095108
Long Paper

A Frequency-Aware Multi-Scale Transformer for Oracle Bone Inscription Semantic Completion

Chao Chen
School of Information Resources Management, Renmin University of China · 2021201248@ruc.edu.cn
Hezi Li
School of Liberal Arts, Renmin University of China · 2022200808@ruc.edu.cn
Zekun Yang
School of Information Resources Management, Renmin University of China; Research Center for Digital Humanities, Renmin University of China · zekunyang@ruc.edu.cn

Introduction

Oracle Bone Inscriptions (OBI) constitute the earliest systematic form of Chinese writing and are indispensable primary sources for understanding Shang dynasty civilization (Qiu 2015). The task of semantic completion, predicting missing or damaged characters from contextual evidence, is fundamental to the philological study of these texts (Mo / Zhang 2023: 47-56; Wang / Cheng 2025: 121-130). While computational methods have made strides in the physical reconstruction of oracle bones through image restoration (Jiang et al. 2024) or fragment matching (Wen et al. 2025: 1-7), they largely neglect the linguistic dimension of the inscriptions. A restoration that focuses solely on physical structure while ignoring contextual linguistic information may lead to unreliable or meaningless results.

The application of modern NLP techniques to this domain is hindered by three fundamental challenges specific to OBI: (1) Extreme data scarcity, with only about two thousand deciphered characters available (Chen 2019: 156-160); (2) Absence of standardized digital encoding due to pervasive graphical variants and undeciphered signs; (3) Long-tail character frequency distribution (Chai 2010; Chen 2007). These characteristics make conventional, data-intensive pre-trained language models ineffective.

This study directly addresses these challenges with two core contributions of significant value to digital humanities and low-resorce NLP. First, we construct the Oracle Bone Inscriptions Corpus (OBIC), a machine-readable dataset specifically engineered for semantic modeling. Second, we propose the Frequency-Aware Multi-scale Transformer (FAMT), a neural architecture that integrates explicit mechanisms to handle data imbalance and capture contextual dependencies at multiple granularities. Our work demonstrates that the semantic completion solution presented in this study achieves state-of-the-art performance, offering humanities scholars a powerful auxiliary tool. This research not only advances the digital study of ancient script corpora, but also provides new methodologies for processing other low-resource, non-standard languages.

The source code and data for this project are openly available on GitHub at https://github.com/aloryori/FAMT.git.

Data

A foundational contribution of this work is the creation of OBIC, a dataset enabling rigorous semantic completion research. Constructed from the comprehensive Oracle Bone Inscriptions Multi-modal Dataset (OBIMD) (Li et al. 2024), OBIC overcomes the primary obstacle of machine readability. We implement a structured, database-inspired mapping scheme: each semantically distinct character concept is assigned a unique primary identifier, and all known graphical variants are linked to this identifier while retaining distinct codes. This two-level schema effectively decouples linguistic meaning from visual form, allowing computational models to learn the equivalence of different shapes representing the same character.

Applying this encoding and filtering out sequences too short or containing too many undeciphered characters, we yield a core corpus of 13,527 sentences, comprising 1,660 distinct characters and 81,084 tokens. Statistical analysis confirms its long-tail distribution (see Figure 1). To directly address this imbalance, we design an oversampling strategy (see Figure 2). This strategy increases the sampling weight of sequences containing low-frequency characters, thereby enriching the training signal for rare but critical lexical items without discarding common patterns. The final augmented OBIC dataset of 17,585 sequences, partitioned into standard training, validation, and test sets, represents a significant infrastructural contribution. The dataset provides an essential, high-quality foundation required for developing and benchmarking advanced semantic models in this field.

Figure 1 Cumulative frequency distribution of characters in OBIC sequences

Figure 2 Sequence Oversampling Weight

Method

Our primary methodological innovation is FAMT, a neural architecture explicitly designed to overcome the unique challenges of OBI semantic completion. The FAMT framework follows a streamlined yet powerful three-stage pipeline: adaptive input masking, frequency-aware contextual encoding, and loss-optimized prediction.

The process begins with adaptive masking, a strategic departure from uniform random masking. We assign higher masking probabilities to low-frequency characters, ensuring that the model is rigorously trained to predict these rare but informative tokens. Simultaneously, we reduce masking at sequence boundaries to preserve crucial syntactic cues that guide prediction of interior elements.

Figure 3 The FAMT Architecture

The core of our approach is the FAMT encoder (see Figure 3). It enhances the standard Transformer with two deeply integrated, synergistic innovations. The first is dynamic frequency awareness. Each token’s representation is augmented with a frequency embedding prior to the attention calculation. This allows the model’s attention mechanism to dynamically adjust its focus based on the inherent predictability and information value associated with a token’s rarity, injecting crucial prior knowledge into the learning process.

The second innovation is a multi-scale attention mechanism coupled with a Frequency-Aware Feed-Forward Network (FA-FFN). We configure different attention heads to operate over varying contextual ranges, from local character collocations to broader thematic connections across the entire inscription. This enables FAMT to simultaneously capture fine-grained syntactic patterns and high-level semantic coherence, which is vital for interpreting OBI’s concise yet dense prose. The FA-FFN further amplifies this capability by explicitly modulating the representation strength of low-frequency tokens in the network’s intermediate layers, actively compensating for their under-representation.

For the final prediction stage, we employ the Adaptive Focal Loss (AFL). This loss function automatically down-weights the contribution of easy-to-classify high-frequency characters to the gradient, while sharpening the focus on challenging predictions associated with rare characters. This optimization strategy directly targets the long-tail problem, steering the model’s learning capacity toward the most difficult and semantically valuable cases.

Analysis

We conduct comprehensive experiments on the OBIC test set to evaluate FAMT against strong and diverse baselines, including CopticBiLSTM (Levine et al. 2024: 61-70), CNN-BiLSTM (Singh / Kamboj 2025: 3900-3910), BERT (Devlin et al. 2019: 4171-4186), and RoBERTa (Duan et al. 2024: 14005-14015). The results unequivocally demonstrate FAMT’s superior performance (see Table 1 and Figure 4). FAMT achieves a state-of-the-art accuracy of 0.8051, significantly outperforming all benchmarks. The stark underperformance of RoBERTa (0.2297) highlights the inadequacy of generic, data-intensive models in this ultra-low-resource setting. Beyond accuracy, FAMT excels in recall-oriented metrics and prediction confidence, indicating its outputs are both accurate and reliable.

Table 1 The Comprehensive Performance of Semantic Completion Models

ModelAccHit@3MRRATP
FAMT0.80510.85620.83730.7715
CopticBiLSTM0.68630.76220.73640.5412
CNN-BiLSTM0.50700.63210.58990.2464
BERT0.66240.75810.72350.5400
RoBERTa0.22970.33570.31230.0928

Figure 4 The Performance of Semantic Completion Models on Data with Different Frequency Distributions

Ablation studies (see Table 2) rigorously validate the contribution of each core component. Removing either frequency awareness (-FA) or multi-scale attention (-M) causes a measurable performance drop. Removing both (-FA-M) results in the most significant degradation, confirming their synergistic effect. Further stratified analysis (see Figure 5) reveals their distinct, complementary roles: the frequency module primarily stabilizes predictions for high/medium-frequency characters by correcting data augmentation bias, while the multi-scale mechanism dramatically improves disambiguation for low-frequency characters by integrating broader contextual semantics.

Table 2 The Ablation Experiment Results of FAMT

ModelAccHit@3MRRATP
FAMT0.80510.85620.83730.7715
-FA0.79700.85460.83410.7497
-M0.79700.85150.83220.7443
-FA-M0.77670.84040.81660.7170

Figure 5 The Results of Ablation Experiments on Data with Different Frequency Distributions

Qualitative case studies (see Table 3 and 4) vividly illustrate that baseline models often default to predicting common but contextually loose characters, whereas FAMT successfully identifies rarer, semantically precise characters that align with the inscription’s overall meaning, demonstrating a qualitatively deeper understanding.

Table 3 A semantic completion case from the comparative experiment

Table 4 A semantic completion case from the ablation experiment

Conclusion

This research provides a systematic computational framework for the semantic completion of Oracle Bone Inscriptions. By constructing the OBIC corpus and proposing the FAMT model architecture, we address three core challenges in this domain: data scarcity, encoding inconsistency, and extreme frequency imbalance. Experimental results demonstrate that our proposed approach achieves state-of-the-art, robust, and reliable performance compared to existing methods.

The value of this work extends significantly beyond oracle bone studies. The methodologies developed, particularly for handling data imbalance through dynamic frequency integration and for modeling multi-scale context in short texts, provide a novel and transferable framework for computational research on other low-resource historical languages and non-standard corpora. This project exemplifies a principled approach to interdisciplinary research, where technical innovations in AI are deeply informed by and directly address concrete, longstanding problems in the humanities, resulting in tools that genuinely augment traditional scholarly practice.

Future work will focus on expanding OBIC with newly deciphered materials, incorporating deeper linguistic annotations, and exploring pathways for integrating such specialized, efficient models with general-purpose large language models. We believe this research direction holds strong promise for bridging the gap between advanced computational techniques and the nuanced demands of historical philology, opening new avenues for interrogating and preserving our written heritage.

References
  1. Chai, Guimin (2010): Research on the Characters of Yinxu Huayuanzhuang Dongdi Jiagu. MA thesis, Soochow University.
  2. Chen, Tingzhu (2007): Research of the Structural System of the Oracle-Bone Inscriptions. Ph.D. thesis, East China Normal University.
  3. Chen, Yingjie (2019): “On the Number of Oracle Bone Script Characters and Related Issues”, in: Chinese Calligraphy 23:156–160.
  4. Devlin, Jacob / Chang, Ming-Wei / Lee, Kenton / Toutanova, Kristina (2019): “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, in: Association for Computational Linguistics (ed.): Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, Minnesota, June 2019: 4171–4186. DOI: 10.18653/v1/N19-1423.
  5. Duan, Siyu / Wang, Jun / Su, Qi (2024): “Restoring Ancient Ideograph: A Multimodal Multitask Neural Network Approach”, in: European Language Resources Association and International Com-mittee on Computational Linguistics (ed.): Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, Torino, Italia, May 2024: 14005–14015.
  6. Jiang, Hanqi / Pan, Yi / Chen, Junhao / Liu, Zhengliang / Zhou, Yifan / Shu, Peng / Li, Yiwei / Zhao, Huaqin / Mihm, Stephen / Howe, Lewis C / Liu, Tianming (2024): “OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion”, in: arXiv preprint, arXiv:2411.17837 < https://arxiv.org/abs/2411.17837> [31.12.2025].
  7. Levine, Lauren / Li, Cindy / Bremer-McCollum, Lydia / Wagner, Nicholas / Zeldes, Amir (2024): “Lacuna Language Learning: Leveraging RNNs for Ranked Text Completion in Digitized Coptic Manuscripts”, in: Association for Computational Linguistics (ed.): Proceedings of the 1st Workshop on Machine Learning for Ancient Languages, Bangkok, Thailand, August 2024: 61-70. DOI: 10.18653/v1/2024.ml4al-1.8.
  8. Li, Bang / Luo, Donghao / Liang, Yujie / Yang, Jing / Ding, Zengmao / Peng, Xu / Jiang, Boyuan / Han, Shengwei / Sui, Dan / Qin, Peichao / Wu, Pian / Wang, Chaoyang / Qi, Yun / Jin, Taisong / Wang, Chengjie / Huang, Xiaoming / Shu, Zhan / Ji, Rongrong / Liu, Yongge / Wu, Yunsheng (2024): “Oracle Bone Inscriptions Multi-modal Dataset”, in: arXiv preprint, arXiv:2407.03900 < https://arxiv.org/abs/2407.03900 > [31.12.2025].
  9. Mo, Bofeng / Zhang, Chongsheng (2023): “The Application and Prospect of Artificial Intelligence in the Study of Paleography”, in: Chinese Culture Research 2: 47–56.
  10. Qiu, Xigui (2015): Collected Academic Works of Qiu Xigui. Shanghai: Fudan University Press.
  11. Singh, Gurpreet / Kamboj, C P (2025): “Next Word Prediction in Social Media Texts for Punjabi-English Bilingual Users with Sequential CNN-BiLSTM”, in: Baghdad Science Journal 22, 11: 3900-3910.
  12. Wang, Fei / Cheng, Bangxiong (2025): “The Characteristics of Dong Zuobin’s Interpretation of Oracle Bone Characters—A Discussion of the Importance of Staging in the Interpretation of the Oracle Bone Script”, in: Studies in Language and Linguistics 45, 1: 121–130.
  13. Wen, Xuan / Setthawong, Rachsuda / Sun, Haimeng (2025): “Oracle Bone Inscriptions Recognition Based on Deep Learning”, in: Institute of Electrical and Electronics Engineers (ed.): Second International Conference on Electronics, Communications and Intelligent Science, Yueyang, China, May 2025: 1–7. DOI: 10.1109/ECIS65594.2025.