DH 2026

Daejeon, July 27–31

Fri, July 3111:00–12:30S041105
Long Paper

An Interpretable Stylometric Framework for Early Chinese Buddhist Translator Styles: A Case Study on the Attribution of Disputed Works of An Shigao

Yating Pan
Department of Information Management, Peking University, Beijing, China; Research Center for Digital Humanities, Peking University, Beijing, China · ytpan25@stu.pku.edu.cn
Xinyue Zhang
Department of Philosophy and Religious Studies, Peking University, Beijing, China · 2501211060@stu.pku.edu.cn
Jun Wang
Department of Information Management, Peking University, Beijing, China; Research Center for Digital Humanities, Peking University, Beijing, China · junwang@pku.edu.cn

Introduction

The history of Chinese Buddhist translation (2nd–10th c.) marks a pivotal stage in cultural exchange among India, Central Asia, and China (Zürcher 1991; Nattier 2008), yet how translators’ stylistic choices evolved remains underexplored. Traditional philological research, from Zürcher (1991) to Hung / Hsieh / Bingenheimer (2017), identified translator-specific traits but lacked systematic quantification. Recent computational work has advanced this field: Hung et al. (2017) applied PCA to trace diachronic stylistic change; Bingenheimer (2020) used text mining for inter-Āgama relations; and Lu et al. (2025) employed BERT for attribution with 95.2% accuracy.

Although deep models like BERT achieve strong results, they remain black boxes for philological argumentation. In contrast, traditional approaches such as stylometry perform well in classical languages (Gianitsos et al. 2019) and remain widely used in digital humanities for their interpretability (Holmes 1998). We therefore adopt a domain-expert-guided, interpretable stylometric framework that extracts linguistically grounded features from CBETA scriptures (CBETA 2025) and statistically validates them. Applied to disputed works of An Shigao, the framework produces results consistent with both philological scholarship (Nattier 2008; Zacchetti 2010) and BERT-based models, demonstrating that interpretable stylometry can rival deep learning while remaining transparent and reproducible, without any supervised training.

Data Collection and Preprocessing

We compile a CBETA-based corpus of authenticated translations by seven early translators from Eastern Han to Western Jin: Yan Fotiao, An Shigao, Kang Senghui, Kang Mengxiang, Lokakṣema, Zhi Qian and Dharmarakṣa. The set comprises 31 texts, providing stylistic baselines for early translation schools (Appendix A, Table A1). A secondary evaluation set includes eight disputed works historically associated with An Shigao in classical catalogues but questioned or reassessed by modern philology (Appendix A, Table A2).

Preprocessing removes titles and annotations, retains CJK characters, segments sentences using CBETA punctuation, and preserves intra-sentence punctuation to maintain sentence-pattern signals. All analyses are performed at the document level with subsequent aggregation by translator and period.

Methods

Figure 1 outlines our expert-in-the-loop pipeline. A Buddhist philology expert defines and validates linguistic feature families, while an NLP engineer extracts these features from authenticated and disputed translations to build author and document style profiles. The classification module then compares them to produce attribution results.

Figure 1. Overview of the expert-in-the-loop stylometry pipeline.

Feature Engineering with Expert Validation

We constructed five interpretable stylometric feature families. Two are expert-defined (function words; sentence-formula templates). Three are data-driven (sentence-final characters; sentence-length; POS n-grams) and expert-verified. Feature details are shown in Appendix B, Tables B1–B4.

Let a document have S sentences (s₁,…,s_S). Define total CJK characters T = Σ |CJK(s_i)| over i = 1…S. Define POS-tag sequence P = (p₁,…,p_L).

Function words: Philology experts defined twenty function words covering pronouns, prepositions, particles, adverbs and conjunctions. For each function word funcw, we computed its frequency per 100 words:

Sentence formulas: Philology experts defined six rhetorical sentence formulas (enumerative, definitional, categorical, interrogative-enumerative, direct-quotation and conditional-temporal) to encode the translators’ personal pedagogical and exegetical expressions. For each formula f, let 1f(s_i) = 1 if sentence s_i matches f once. A sentence may match multiple templates independently, so we computed each pattern’s proportion out of the total sentence count:

Sentence-final characters: We mined the top-10 sentence-final characters that most frequently appear across more than half of the translators to reflect their habitual closure styles. For each character c, we computed its frequency among total sentence counts:

Sentence-length distribution: We computed the mean, median and standard deviation of sentence length for each document to quantify a translator’s structural rhythm and rhetorical density. Short, uniform sentences signal segmented oral prose, whereas long and connective elaborations indicate a shift toward literati styles.

POS n-grams: Part-of-speech n-grams capture stylistic consistency in local syntactic dependencies. We used the Jiayan Toolkit to build a global vocabulary of top-k POS n-grams (n = 1..4) from the corpus. For each n, we computed the relative frequency of every vocabulary n-gram pattern g (e.g., g = nv, g = vvu, etc.) within a document’s POS-tag sequence P = (p₁,…,p_L):

Evidence-Enhanced Translator Styling

We grouped all features by translator and applied the Kruskal–Wallis (H) test with Holm–Bonferroni correction (p_holm < 0.01). The ten most discriminative features (Appendix B, Table B4) were retained, and standardised z-scores

were computed to build normalised stylistic fingerprints. Translator profiles were then mapped to 0–100 indices via corpus-wise z-scaling and sigmoid transformation, with pooled standard deviations stabilising small-sample variation. Results across translations and translators are in Appendix C, Tables C1 and C2.

Attribution Scoring for Disputed Translations

Each disputed text was transformed into the same sigmoid-normalised space as the author profiles by z-scoring each feature (using global μ and σ from non-disputed texts) and applying a sigmoid. Similarity was then computed as the inverse of a standardised chi-square distance across features, yielding Sim ∈ [0, 1]. We interpret Sim comparatively, with Sim ≥ 0.60 indicating high confidence.

Results and Discussion

Translator Stylometric Differentiation

Figure 2 visualises the ten core stylometric features across the seven translators stratified by historical period, revealing the interconnection between individual translator style and Buddhist translation norms. The earliest Eastern Han translators, An Shigao and Yan Fotiao, exemplify literal, plain and colloquial translation, marked by extremely low function-word density and minimal 4-grams. Notably, An Shigao’s many long sentences with over 40 characters contain no function words (T0031: 以应念法不念,不应念法者便念,便爱流生,已爱流生便增多不致,未生爱流亦痴流便生,已生爱流亦痴流便增多不致). Instead, his Verb–Adverb 2-gram (70.7%) and Adverb–Verb–Adverb 3-gram scores (71.9%) dominate, exemplified by the exhortatory “为+不+verb” pattern (T0014: 一为不恭敬佛、二为不恭敬法、三为不恭敬同学者、四为不恭敬戒), reflecting the Eastern Han’s oral transmission tradition.

In contrast, Lokakṣema’s function words increase (之: 44.7%; 其: 75.1%), signalling a shift toward formal expression. By the Three Kingdoms period, translators like Zhi Qian and Kang Senghui show systematic growth in connective density (之: 66.4% and 87.6%), replacing literal translation with natural cohesion through complex nominal modification (Zhi Qian T0054: 有淫之意、有怒之意、有痴之意; Kang Senghui T0152: 今又犯之,种无量之罪).

Dharmarakṣa represents the apex, displaying a mature written style. His elevated function-word integration (之: 72.6%; 及: 81.8%) and higher-order structures (pos4_n_v_r_n: 71.5%) exemplify Noun–Verb–Pronoun–Noun subordination (T0154: 菩萨入之心,释子告其人), demonstrating Buddhist sūtra integration into China’s literary tradition. This evidence traces the sinicisation trajectory from the Eastern Han’s plain vernacular to the Western Jin’s refined written expression (Hung et al. 2017).

Figure 2. Stylometric radar charts of early Chinese Buddhist translators based on the top-10 features.

Case Study: An Shigao’s Disputed Translation Attribution

Figure 3. Stylistic space of early Chinese Buddhist translations in multidimensional space.

Figure 4. Per-feature similarity distribution between disputed translations and An Shigao’s stylometric style.

Our attribution results reveal two clear stylistic groups among the disputed translations. T1557 (Abhidharma Prakaraṇa Sūtra) and T0101 (Saṃyuktāgama Sūtra) cluster closely with the verified An Shigao corpus, whereas the other five texts diverge sharply. As shown in Figure 3, the red ellipse marks the centroid and 95% confidence interval of authentic An Shigao works; blue dots denote other translators from the Eastern Han to Western Jin; yellow stars represent disputed texts. T1557 lies at the centre of the An Shigao cluster (Sim = 0.783), displaying his hallmark concise syntax, balanced sentence-formula use and restrained function-word frequency, consistent with Nattier (2008) and Zacchetti (2010). T0101 (Sim = 0.687) also falls within this range, sharing the rhythmic compactness and pedagogical tone described by Lin (2001), Harrison (2002) and Hayashiya (1941). In contrast, T0109, T0105, T0605, T0792 and T0604 lie far outside the An Shigao region (Sim < 0.3), featuring poetic passages and lexical registers characteristic of other or even later translators (Nattier 2008; Zürcher 1991; Hayashiya 1941).

Figure 4 visualises feature-level similarity across the ten core stylometric indicators. Both T1557 and T0101 show high correspondence with An Shigao’s function-word ratios and 4-gram density, providing linguistic validation of their authenticity. These findings also align with Lu et al. (2025)’s BERT-based attribution model, yet our zero-training, expert-interpretable framework offers transparent, feature-based explanations suitable for philological inquiry.

Conclusion

We presented a domain-expert-guided, fully interpretable stylometric framework for early Chinese Buddhist translation studies. Using expert-guided linguistic features and stylometric statistics, the model reconstructs stylistic evolution across the Eastern Han–Western Jin and attributes disputed texts without task-specific training. In the case of An Shigao’s disputed translation attribution, our results match the advanced deep-learning method results. Unlike neural classifiers that require labelled training and offer limited rationale, our pipeline attains comparable results while returning human-interpretable evidence.

We note some practical limitations affecting the interpretation. Shorter texts (e.g., T0604, T0605) have fewer features and lower feature stability; the current focus is mainly on early Chinese translations of Buddhist scriptures. We will expand to longer corpora and later translators (e.g., Kumārajīva and Xuanzang from the Tang Dynasty), and publish a catalogue associated with the evidence to facilitate reproduction. Combining our interpretation with neural models is a promising direction for hybrid, interpretable digital humanities.

Appendix A. Dataset Composition

Table A1. Authenticated translations and their metadata (CBETA).

TranslatorTitle (English)IDVolumesSentencesCharacters
Yan FotiaoSix Pāramitās Sūtra T07781521069
An ShigaoLong Āgama on Ten DharmasT0013257410090
Sūtra on the Desire for Human LifeT001415296255
Sūtra on the Four Noble TruthsT003111011858
Sūtra on the Roots of SufferingT003212393655
Sūtra on the Four Truths of SufferingT0036132729
Sūtra on the True DharmaT00481441330
Discourse on GivingT005711262760
Prajñā Sūtra T009811253576
Eightfold Path SūtraT0112144626
Seven Contemplations SūtraT0150A160210472
Nine Reflections SūtraT0150B117489
Mindfulness SūtraT060323095772
Earth SūtraT060711986529
Kang SenghuiSūtra on the Six PāramitāsT01528330856217
Kang MengxiangSūtra on the Four Tours of the BuddhaT01371581166
Sūtra on the Rise of the BuddhaT0197266113537
LokakṣemaSūtra of the Bodhisattva PathT022410277358944
Sūtra on Ghosts and SpiritsT02801682145
Sūtra on the Samādhi of the VesselT041713176415
Zhi QianSūtra of Four SonsT00541751651
Sūtra of the Brahmin and the DemonsT006812164617
Sūtra on the Transformation of MeritT007611413965
Sūtra of Prince Sudhana’s AwakeningT018524768370
Sūtra of the Bodhisattva’s KarmaT028111945285
Sūtra of VaiśravaṇaT04742113722930
DharmarakṣaSūtra of LifeT01545239848319
Sūtra on Universal LightT01868213658907
Sūtra on Radiant LightT022210387282324
Sūtra on the True DharmaT026310283376382
Mahāparinirvāṇa Sūtra T03988137938580

Table A2. Eight disputed translations attributed to An Shigao and their metadata (CBETA).

Title (English)CBETA IDVolumesSentencesCharacters
Saṃyuktāgama SūtraT010114289247
Sūtra on the Five AggregatesT0105151787
Dhammacakkappavattana SuttaT0109132770
Thirty-Seven Factors of EnlightenmentT0604116931
Sūtra on Meditation and ContemplationT0605110281
Sūtra on Receiving the DharmaT0792117304
Abhidharma Prakaraṇa Sūtra T155713854368
Mahāsaṃnipāta Sūtra, Chapter on Ten-Direction Bodhisattvas T0397_05913169809

Appendix B. Feature System for Stylometric Analysis

Table B1. Sentence-final character features.

Feature TypeCharacters (Chinese)Transliteration
Sentence-final char.也, 是, 之, 乎, 者, 故, 何, 行, 提, 法ye, shi, zhi, hu, zhe, gu, he, xing, ti, fa

Table B2. Sentence-formula features.

TemplateExample (Chinese)Function (English)
第 X第一、第二…Sequential enumeration
X 者色者、心者…Definition or exemplification
以 X 为 Y以心为界、以苦为道Categorical relation
何等故举何等、何者、云何Interrogative / doctrinal question
佛言佛言:…Quotation formula
已 X 便已灭便得道Temporal / conditional sequence

Table B3. Function-character features.

CategoryFunction Words (Chinese)English Gloss
Pronouns之, 其, 自, 皆, 是pronouns (of, its, self, all, to be)
Prepositions以, 於, 为, 与, 及prepositions (by, in, for, with, and)
Particles所, 者, 故particles (that which, one who, therefore)
Conjunctions / Adverbs而, 有, 则, 无, 不, 如, 能conjunctions / adverbs (and, have, then, none, not, like, can)

Table B4. Top-10 most discriminative stylometric features across translators (Kruskal–Wallis p < 0.01).

FeatureDescriptionp_holmε²
func_之_per100Frequency of particle zhi (之, nominaliser) per 100 chars4.80 × 10⁻²0.838
func_其_per100Frequency of pronoun qi (其, ‘his/her/its’) per 100 chars7.22 × 10⁻²0.798
pos3_v_u_n_propPOS trigram Verb–Auxiliary–Noun proportion1.19 × 10⁻¹0.749
func_皆_per100Frequency of adverb jie (皆, ‘all’) per 100 chars1.33 × 10⁻¹0.738
func_及_per100Frequency of conjunction ji (及, ‘and’) per 100 chars1.36 × 10⁻¹0.735
pos3_d_v_d_propPOS trigram Adverb–Verb–Adverb proportion1.52 × 10⁻¹0.724
pos3_n_u_v_propPOS trigram Noun–Auxiliary–Verb proportion1.66 × 10⁻¹0.715
pos4_n_v_r_n_propPOS 4-gram Noun–Verb–Particle–Noun proportion1.88 × 10⁻¹0.702
pos2_v_d_propPOS bigram Verb–Adverb proportion1.88 × 10⁻¹0.702
pos2_n_r_propPOS bigram Noun–Particle proportion2.02 × 10⁻¹0.694

Note: p_holm = Holm–Bonferroni-corrected significance level; ε² = effect size representing the proportion of variance in feature ranks explained by translator differences.

Appendix C. Results Details

Table C1. Stylometric score distribution across each scripture translation.

AuthorScripturefunc_之func_其func_皆func_及pos2_v_dpos2_n_rpos3_v_u_npos3_d_v_dpos3_n_u_vpos4_n_v_r_n
Yan FotiaoT077841.10%28.47%46.27%29.34%37.14%31.10%50.81%39.18%24.77%23.18%
An ShigaoT001329.07%28.47%31.49%31.06%71.76%28.34%30.59%74.50%23.47%27.04%
T001429.07%28.47%31.54%29.34%56.83%37.28%27.95%63.44%30.22%34.96%
T003129.07%28.47%30.85%29.34%89.42%47.01%22.98%94.63%29.36%56.40%
T003229.07%29.46%32.40%29.34%82.59%23.21%26.95%72.50%27.32%26.67%
T003629.07%28.47%30.85%29.34%66.03%20.82%22.98%60.89%17.11%23.18%
T004829.07%28.47%30.85%29.34%56.20%32.37%22.98%75.20%17.11%23.18%
T005729.07%28.47%30.85%29.34%92.12%20.82%28.62%78.34%19.60%23.18%
T009829.07%28.47%33.84%29.34%77.20%22.05%36.72%74.74%27.83%26.83%
T011229.07%28.47%39.69%29.34%62.24%60.91%37.12%77.30%49.02%50.09%
T0150A30.03%29.64%33.33%30.98%57.49%35.65%32.35%59.69%44.95%32.87%
T0150B29.07%28.47%30.85%29.34%90.58%20.82%22.98%79.70%34.60%23.18%
T060329.69%28.47%33.26%29.34%75.66%28.88%30.77%76.69%25.87%35.11%
T060729.07%28.47%36.63%40.01%40.94%34.02%26.57%46.82%38.03%29.76%
LokakṣemaT022446.09%57.73%51.25%59.22%40.15%64.21%93.69%43.25%56.12%42.04%
T028031.94%94.08%98.56%45.29%27.38%97.86%39.66%32.28%90.10%54.77%
T041756.05%73.50%51.40%76.73%44.03%65.97%56.31%53.19%66.19%73.55%
Zhi QianT005461.33%57.31%87.32%29.34%31.44%49.85%51.59%54.16%60.63%81.60%
T006861.66%52.03%68.49%55.16%35.25%54.11%63.11%38.49%57.14%52.36%
T007692.93%70.20%57.07%29.34%30.63%36.31%78.03%29.91%84.56%57.84%
T018563.69%65.90%60.22%64.56%36.08%56.95%67.89%38.31%62.58%60.64%
T028156.21%38.91%62.61%58.18%24.32%51.56%58.79%16.64%63.59%41.83%
T047462.34%69.91%59.78%60.86%39.22%69.31%64.37%41.67%58.66%68.66%
Kang SenghuiT015287.55%68.88%42.92%41.92%31.83%53.82%86.78%28.18%58.39%73.03%
Kang MengxiangT013736.04%56.41%37.47%82.46%25.33%65.30%48.04%15.28%90.45%94.17%
T019747.94%45.09%65.52%87.66%29.00%57.32%41.18%30.61%67.66%74.63%
DharmarakṣaT015477.08%73.08%53.53%80.07%35.28%66.23%60.52%33.17%54.90%80.14%
T018670.49%77.27%73.62%86.61%29.95%67.92%51.31%28.54%59.57%74.02%
T022256.91%57.02%43.77%75.18%45.74%72.51%90.70%41.77%55.97%62.38%
T026380.18%78.50%71.62%89.11%30.68%69.20%58.98%28.58%72.41%74.82%
T039878.41%89.28%53.07%77.98%35.01%70.26%57.20%29.93%77.29%66.11%

Table C2. Aggregated stylometric score distribution across each translator.

Authorfunc_之func_其func_皆func_及pos2_v_dpos2_n_rpos3_v_u_npos3_d_v_dpos3_n_u_vpos4_n_v_r_n
Yan Fotiao41.10%28.47%46.27%29.34%37.14%31.10%50.81%39.18%24.77%23.18%
An Shigao29.19%28.64%32.80%30.42%70.70%31.71%28.43%71.88%29.58%31.73%
Kang Senghui87.55%68.88%42.92%41.92%31.83%53.82%86.78%28.18%58.39%73.03%
Kang Mengxiang41.99%50.75%51.49%85.06%27.17%61.31%44.61%22.94%79.05%84.40%
Lokakṣema44.70%75.10%67.07%60.41%37.19%76.01%63.22%42.91%70.80%56.79%
Zhi Qian66.36%59.04%65.91%49.57%32.82%53.02%63.96%36.53%64.53%60.49%
Dharmarakṣa72.61%75.03%59.12%81.79%35.33%69.22%63.74%32.40%64.03%71.49%

Table C3. Classification of An Shigao’s disputed translations.

CBETA IDOur ResultBERT Model Result (Lu et al. 2025)
T0101YesYes
T0105NoNo
T0109NoNo
T0604NoNo
T0605NoNo
T0792NoNo
T1557An ShigaoAn Shigao
T0397_059NoNo
References
  1. Bingenheimer, Marcus (2020): “A Study and Translation of the Yakṣa-saṃyukta in the Shorter Chinese Saṃyukta-āgama”, in: Dhammadinnā (ed.): Research on the Saṃyukta-āgama. Taipei: Dharma Drum Publishing Corporation, 763–841. ISBN 978-957-598-859-3.
  2. CBETA (2025): Chinese Buddhist Electronic Text Association (CBETA) digital corpus [Dataset]. <https://cbeta.org> [03.05.2026].
  3. Gianitsos, Efthimios Tim / Bolt, Thomas J. / Chaudhuri, Pramit / Dexter, Joseph P. (2019): “Stylometric Classification of Ancient Greek Literary Texts by Genre”, in: Proceedings of the 3rd Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2019). Stroudsburg, PA: Association for Computational Linguistics, 52–60. DOI: 10.18653/v1/W19-2507.
  4. Harrison, Paul (2002): “Another Addition to the An Shigao Corpus? Preliminary Notes on an Early Chinese Saṃyukta-āgama Translation”, in: Early Buddhism and Abhidharma Thought: In Honour of Dr Hajime Sakurabe on His Seventy-seventh Birthday. Kyoto: Heirakuji Shoten, 1–32.
  5. Hayashiya, Tomojirō 林屋友次郎 (1941): Kyōroku kenkyū (zenpen) 經録研究(前篇). Tokyo: Iwanami Shoten.
  6. Holmes, David I. (1998): “The Evolution of Stylometry in Humanities Scholarship”, in: Literary and Linguistic Computing 13, 3: 111–117. DOI: 10.1093/llc/13.3.111.
  7. Hung, Jen-jou / Hsieh, Cheng-en / Bingenheimer, Marcus (2017): “Stylometric Analysis of Chinese Buddhist Texts — Do Different Chinese Translations of the Gaṇḍavyūha Reflect Stylistic Features That Are Typical for Their Age?”, in: Journal of the Japanese Association for Digital Humanities 2, 1: 1–30. DOI: 10.17928/jjadh.2.1_1.
  8. Jiayan Toolkit: Jiayan: NLP Toolkit for Classical Chinese [Software]. GitHub. <https://github.com/jiaeyan/Jiayan> [03.05.2026].
  9. Lin, Yueh-Mei (2001): A Study on the Anthology Za Ahan Jing (T101) Centred on Its Linguistic Features, Authorship and School Affiliation. MA thesis, University of Canterbury. <https://hdl.handle.net/10092/104838> [03.05.2026].
  10. Lu, Lu / Zheng, Yi / Fang, Yixin (2025): “Utilizing Language Models for the Attribution of Chinese Buddhist Translations: A Case Study of An Shigao’s Works” (以安世高译经考辨为例), in: Journal of Zhejiang University (Humanities and Social Sciences) 55, 2. DOI: 10.3785/j.issn.1008-942X.CN33-6000/C.2024.03.061.
  11. Nattier, Jan (2008): A Guide to the Earliest Chinese Buddhist Translations: Texts from the Eastern Han and Three Kingdoms Periods. Tokyo: IRIAB, Soka University. ISBN 978-4-904234-00-6.
  12. Zacchetti, Stefano (2010): “Defining An Shigao’s Translation Corpus: The State of the Art in Relevant Research”, in: Shen, Wei-Rong (ed.): Historical and Philological Studies of China’s Western Regions, No. 3. Beijing: Kexue Chubanshe, 249–270.
  13. Zürcher, Erik (1991): “A New Look at the Earliest Chinese Buddhist Texts”, in: Shinohara, Koichi / Schopen, Gregory (eds.): From Benares to Beijing: Essays on Buddhism and Chinese Religion in Honour of Prof. Jan Yün-hua. Oakville, ON: Mosaic Press, 277–304. ISBN 978-0-88962-444-3.