Daejeon, July 27–31
The history of Chinese Buddhist translation (2nd–10th c.) marks a pivotal stage in cultural exchange among India, Central Asia, and China (Zürcher 1991; Nattier 2008), yet how translators’ stylistic choices evolved remains underexplored. Traditional philological research, from Zürcher (1991) to Hung / Hsieh / Bingenheimer (2017), identified translator-specific traits but lacked systematic quantification. Recent computational work has advanced this field: Hung et al. (2017) applied PCA to trace diachronic stylistic change; Bingenheimer (2020) used text mining for inter-Āgama relations; and Lu et al. (2025) employed BERT for attribution with 95.2% accuracy.
Although deep models like BERT achieve strong results, they remain black boxes for philological argumentation. In contrast, traditional approaches such as stylometry perform well in classical languages (Gianitsos et al. 2019) and remain widely used in digital humanities for their interpretability (Holmes 1998). We therefore adopt a domain-expert-guided, interpretable stylometric framework that extracts linguistically grounded features from CBETA scriptures (CBETA 2025) and statistically validates them. Applied to disputed works of An Shigao, the framework produces results consistent with both philological scholarship (Nattier 2008; Zacchetti 2010) and BERT-based models, demonstrating that interpretable stylometry can rival deep learning while remaining transparent and reproducible, without any supervised training.
We compile a CBETA-based corpus of authenticated translations by seven early translators from Eastern Han to Western Jin: Yan Fotiao, An Shigao, Kang Senghui, Kang Mengxiang, Lokakṣema, Zhi Qian and Dharmarakṣa. The set comprises 31 texts, providing stylistic baselines for early translation schools (Appendix A, Table A1). A secondary evaluation set includes eight disputed works historically associated with An Shigao in classical catalogues but questioned or reassessed by modern philology (Appendix A, Table A2).
Preprocessing removes titles and annotations, retains CJK characters, segments sentences using CBETA punctuation, and preserves intra-sentence punctuation to maintain sentence-pattern signals. All analyses are performed at the document level with subsequent aggregation by translator and period.
Figure 1 outlines our expert-in-the-loop pipeline. A Buddhist philology expert defines and validates linguistic feature families, while an NLP engineer extracts these features from authenticated and disputed translations to build author and document style profiles. The classification module then compares them to produce attribution results.
Figure 1. Overview of the expert-in-the-loop stylometry pipeline.
We constructed five interpretable stylometric feature families. Two are expert-defined (function words; sentence-formula templates). Three are data-driven (sentence-final characters; sentence-length; POS n-grams) and expert-verified. Feature details are shown in Appendix B, Tables B1–B4.
Let a document have S sentences (s₁,…,s_S). Define total CJK characters T = Σ |CJK(s_i)| over i = 1…S. Define POS-tag sequence P = (p₁,…,p_L).
Function words: Philology experts defined twenty function words covering pronouns, prepositions, particles, adverbs and conjunctions. For each function word funcw, we computed its frequency per 100 words:
Sentence formulas: Philology experts defined six rhetorical sentence formulas (enumerative, definitional, categorical, interrogative-enumerative, direct-quotation and conditional-temporal) to encode the translators’ personal pedagogical and exegetical expressions. For each formula f, let 1f(s_i) = 1 if sentence s_i matches f once. A sentence may match multiple templates independently, so we computed each pattern’s proportion out of the total sentence count:
Sentence-final characters: We mined the top-10 sentence-final characters that most frequently appear across more than half of the translators to reflect their habitual closure styles. For each character c, we computed its frequency among total sentence counts:
Sentence-length distribution: We computed the mean, median and standard deviation of sentence length for each document to quantify a translator’s structural rhythm and rhetorical density. Short, uniform sentences signal segmented oral prose, whereas long and connective elaborations indicate a shift toward literati styles.
POS n-grams: Part-of-speech n-grams capture stylistic consistency in local syntactic dependencies. We used the Jiayan Toolkit to build a global vocabulary of top-k POS n-grams (n = 1..4) from the corpus. For each n, we computed the relative frequency of every vocabulary n-gram pattern g (e.g., g = nv, g = vvu, etc.) within a document’s POS-tag sequence P = (p₁,…,p_L):
We grouped all features by translator and applied the Kruskal–Wallis (H) test with Holm–Bonferroni correction (p_holm < 0.01). The ten most discriminative features (Appendix B, Table B4) were retained, and standardised z-scores
were computed to build normalised stylistic fingerprints. Translator profiles were then mapped to 0–100 indices via corpus-wise z-scaling and sigmoid transformation, with pooled standard deviations stabilising small-sample variation. Results across translations and translators are in Appendix C, Tables C1 and C2.
Each disputed text was transformed into the same sigmoid-normalised space as the author profiles by z-scoring each feature (using global μ and σ from non-disputed texts) and applying a sigmoid. Similarity was then computed as the inverse of a standardised chi-square distance across features, yielding Sim ∈ [0, 1]. We interpret Sim comparatively, with Sim ≥ 0.60 indicating high confidence.
Figure 2 visualises the ten core stylometric features across the seven translators stratified by historical period, revealing the interconnection between individual translator style and Buddhist translation norms. The earliest Eastern Han translators, An Shigao and Yan Fotiao, exemplify literal, plain and colloquial translation, marked by extremely low function-word density and minimal 4-grams. Notably, An Shigao’s many long sentences with over 40 characters contain no function words (T0031: 以应念法不念,不应念法者便念,便爱流生,已爱流生便增多不致,未生爱流亦痴流便生,已生爱流亦痴流便增多不致). Instead, his Verb–Adverb 2-gram (70.7%) and Adverb–Verb–Adverb 3-gram scores (71.9%) dominate, exemplified by the exhortatory “为+不+verb” pattern (T0014: 一为不恭敬佛、二为不恭敬法、三为不恭敬同学者、四为不恭敬戒), reflecting the Eastern Han’s oral transmission tradition.
In contrast, Lokakṣema’s function words increase (之: 44.7%; 其: 75.1%), signalling a shift toward formal expression. By the Three Kingdoms period, translators like Zhi Qian and Kang Senghui show systematic growth in connective density (之: 66.4% and 87.6%), replacing literal translation with natural cohesion through complex nominal modification (Zhi Qian T0054: 有淫之意、有怒之意、有痴之意; Kang Senghui T0152: 今又犯之,种无量之罪).
Dharmarakṣa represents the apex, displaying a mature written style. His elevated function-word integration (之: 72.6%; 及: 81.8%) and higher-order structures (pos4_n_v_r_n: 71.5%) exemplify Noun–Verb–Pronoun–Noun subordination (T0154: 菩萨入之心,释子告其人), demonstrating Buddhist sūtra integration into China’s literary tradition. This evidence traces the sinicisation trajectory from the Eastern Han’s plain vernacular to the Western Jin’s refined written expression (Hung et al. 2017).
Figure 2. Stylometric radar charts of early Chinese Buddhist translators based on the top-10 features.
Figure 3. Stylistic space of early Chinese Buddhist translations in multidimensional space.
Figure 4. Per-feature similarity distribution between disputed translations and An Shigao’s stylometric style.
Our attribution results reveal two clear stylistic groups among the disputed translations. T1557 (Abhidharma Prakaraṇa Sūtra) and T0101 (Saṃyuktāgama Sūtra) cluster closely with the verified An Shigao corpus, whereas the other five texts diverge sharply. As shown in Figure 3, the red ellipse marks the centroid and 95% confidence interval of authentic An Shigao works; blue dots denote other translators from the Eastern Han to Western Jin; yellow stars represent disputed texts. T1557 lies at the centre of the An Shigao cluster (Sim = 0.783), displaying his hallmark concise syntax, balanced sentence-formula use and restrained function-word frequency, consistent with Nattier (2008) and Zacchetti (2010). T0101 (Sim = 0.687) also falls within this range, sharing the rhythmic compactness and pedagogical tone described by Lin (2001), Harrison (2002) and Hayashiya (1941). In contrast, T0109, T0105, T0605, T0792 and T0604 lie far outside the An Shigao region (Sim < 0.3), featuring poetic passages and lexical registers characteristic of other or even later translators (Nattier 2008; Zürcher 1991; Hayashiya 1941).
Figure 4 visualises feature-level similarity across the ten core stylometric indicators. Both T1557 and T0101 show high correspondence with An Shigao’s function-word ratios and 4-gram density, providing linguistic validation of their authenticity. These findings also align with Lu et al. (2025)’s BERT-based attribution model, yet our zero-training, expert-interpretable framework offers transparent, feature-based explanations suitable for philological inquiry.
We presented a domain-expert-guided, fully interpretable stylometric framework for early Chinese Buddhist translation studies. Using expert-guided linguistic features and stylometric statistics, the model reconstructs stylistic evolution across the Eastern Han–Western Jin and attributes disputed texts without task-specific training. In the case of An Shigao’s disputed translation attribution, our results match the advanced deep-learning method results. Unlike neural classifiers that require labelled training and offer limited rationale, our pipeline attains comparable results while returning human-interpretable evidence.
We note some practical limitations affecting the interpretation. Shorter texts (e.g., T0604, T0605) have fewer features and lower feature stability; the current focus is mainly on early Chinese translations of Buddhist scriptures. We will expand to longer corpora and later translators (e.g., Kumārajīva and Xuanzang from the Tang Dynasty), and publish a catalogue associated with the evidence to facilitate reproduction. Combining our interpretation with neural models is a promising direction for hybrid, interpretable digital humanities.
Table A1. Authenticated translations and their metadata (CBETA).
| Translator | Title (English) | ID | Volumes | Sentences | Characters |
| Yan Fotiao | Six Pāramitās Sūtra | T0778 | 1 | 52 | 1069 |
| An Shigao | Long Āgama on Ten Dharmas | T0013 | 2 | 574 | 10090 |
| Sūtra on the Desire for Human Life | T0014 | 1 | 529 | 6255 | |
| Sūtra on the Four Noble Truths | T0031 | 1 | 101 | 1858 | |
| Sūtra on the Roots of Suffering | T0032 | 1 | 239 | 3655 | |
| Sūtra on the Four Truths of Suffering | T0036 | 1 | 32 | 729 | |
| Sūtra on the True Dharma | T0048 | 1 | 44 | 1330 | |
| Discourse on Giving | T0057 | 1 | 126 | 2760 | |
| Prajñā Sūtra | T0098 | 1 | 125 | 3576 | |
| Eightfold Path Sūtra | T0112 | 1 | 44 | 626 | |
| Seven Contemplations Sūtra | T0150A | 1 | 602 | 10472 | |
| Nine Reflections Sūtra | T0150B | 1 | 17 | 489 | |
| Mindfulness Sūtra | T0603 | 2 | 309 | 5772 | |
| Earth Sūtra | T0607 | 1 | 198 | 6529 | |
| Kang Senghui | Sūtra on the Six Pāramitās | T0152 | 8 | 3308 | 56217 |
| Kang Mengxiang | Sūtra on the Four Tours of the Buddha | T0137 | 1 | 58 | 1166 |
| Sūtra on the Rise of the Buddha | T0197 | 2 | 661 | 13537 | |
| Lokakṣema | Sūtra of the Bodhisattva Path | T0224 | 10 | 2773 | 58944 |
| Sūtra on Ghosts and Spirits | T0280 | 1 | 68 | 2145 | |
| Sūtra on the Samādhi of the Vessel | T0417 | 1 | 317 | 6415 | |
| Zhi Qian | Sūtra of Four Sons | T0054 | 1 | 75 | 1651 |
| Sūtra of the Brahmin and the Demons | T0068 | 1 | 216 | 4617 | |
| Sūtra on the Transformation of Merit | T0076 | 1 | 141 | 3965 | |
| Sūtra of Prince Sudhana’s Awakening | T0185 | 2 | 476 | 8370 | |
| Sūtra of the Bodhisattva’s Karma | T0281 | 1 | 194 | 5285 | |
| Sūtra of Vaiśravaṇa | T0474 | 2 | 1137 | 22930 | |
| Dharmarakṣa | Sūtra of Life | T0154 | 5 | 2398 | 48319 |
| Sūtra on Universal Light | T0186 | 8 | 2136 | 58907 | |
| Sūtra on Radiant Light | T0222 | 10 | 3872 | 82324 | |
| Sūtra on the True Dharma | T0263 | 10 | 2833 | 76382 | |
| Mahāparinirvāṇa Sūtra | T0398 | 8 | 1379 | 38580 |
Table A2. Eight disputed translations attributed to An Shigao and their metadata (CBETA).
| Title (English) | CBETA ID | Volumes | Sentences | Characters |
| Saṃyuktāgama Sūtra | T0101 | 1 | 428 | 9247 |
| Sūtra on the Five Aggregates | T0105 | 1 | 51 | 787 |
| Dhammacakkappavattana Sutta | T0109 | 1 | 32 | 770 |
| Thirty-Seven Factors of Enlightenment | T0604 | 1 | 16 | 931 |
| Sūtra on Meditation and Contemplation | T0605 | 1 | 10 | 281 |
| Sūtra on Receiving the Dharma | T0792 | 1 | 17 | 304 |
| Abhidharma Prakaraṇa Sūtra | T1557 | 1 | 385 | 4368 |
| Mahāsaṃnipāta Sūtra, Chapter on Ten-Direction Bodhisattvas | T0397_059 | 1 | 316 | 9809 |
Table B1. Sentence-final character features.
| Feature Type | Characters (Chinese) | Transliteration |
| Sentence-final char. | 也, 是, 之, 乎, 者, 故, 何, 行, 提, 法 | ye, shi, zhi, hu, zhe, gu, he, xing, ti, fa |
Table B2. Sentence-formula features.
| Template | Example (Chinese) | Function (English) |
| 第 X | 第一、第二… | Sequential enumeration |
| X 者 | 色者、心者… | Definition or exemplification |
| 以 X 为 Y | 以心为界、以苦为道 | Categorical relation |
| 何等故举 | 何等、何者、云何 | Interrogative / doctrinal question |
| 佛言 | 佛言:… | Quotation formula |
| 已 X 便 | 已灭便得道 | Temporal / conditional sequence |
Table B3. Function-character features.
| Category | Function Words (Chinese) | English Gloss |
| Pronouns | 之, 其, 自, 皆, 是 | pronouns (of, its, self, all, to be) |
| Prepositions | 以, 於, 为, 与, 及 | prepositions (by, in, for, with, and) |
| Particles | 所, 者, 故 | particles (that which, one who, therefore) |
| Conjunctions / Adverbs | 而, 有, 则, 无, 不, 如, 能 | conjunctions / adverbs (and, have, then, none, not, like, can) |
Table B4. Top-10 most discriminative stylometric features across translators (Kruskal–Wallis p < 0.01).
| Feature | Description | p_holm | ε² |
| func_之_per100 | Frequency of particle zhi (之, nominaliser) per 100 chars | 4.80 × 10⁻² | 0.838 |
| func_其_per100 | Frequency of pronoun qi (其, ‘his/her/its’) per 100 chars | 7.22 × 10⁻² | 0.798 |
| pos3_v_u_n_prop | POS trigram Verb–Auxiliary–Noun proportion | 1.19 × 10⁻¹ | 0.749 |
| func_皆_per100 | Frequency of adverb jie (皆, ‘all’) per 100 chars | 1.33 × 10⁻¹ | 0.738 |
| func_及_per100 | Frequency of conjunction ji (及, ‘and’) per 100 chars | 1.36 × 10⁻¹ | 0.735 |
| pos3_d_v_d_prop | POS trigram Adverb–Verb–Adverb proportion | 1.52 × 10⁻¹ | 0.724 |
| pos3_n_u_v_prop | POS trigram Noun–Auxiliary–Verb proportion | 1.66 × 10⁻¹ | 0.715 |
| pos4_n_v_r_n_prop | POS 4-gram Noun–Verb–Particle–Noun proportion | 1.88 × 10⁻¹ | 0.702 |
| pos2_v_d_prop | POS bigram Verb–Adverb proportion | 1.88 × 10⁻¹ | 0.702 |
| pos2_n_r_prop | POS bigram Noun–Particle proportion | 2.02 × 10⁻¹ | 0.694 |
Note: p_holm = Holm–Bonferroni-corrected significance level; ε² = effect size representing the proportion of variance in feature ranks explained by translator differences.
Table C1. Stylometric score distribution across each scripture translation.
| Author | Scripture | func_之 | func_其 | func_皆 | func_及 | pos2_v_d | pos2_n_r | pos3_v_u_n | pos3_d_v_d | pos3_n_u_v | pos4_n_v_r_n |
| Yan Fotiao | T0778 | 41.10% | 28.47% | 46.27% | 29.34% | 37.14% | 31.10% | 50.81% | 39.18% | 24.77% | 23.18% |
| An Shigao | T0013 | 29.07% | 28.47% | 31.49% | 31.06% | 71.76% | 28.34% | 30.59% | 74.50% | 23.47% | 27.04% |
| T0014 | 29.07% | 28.47% | 31.54% | 29.34% | 56.83% | 37.28% | 27.95% | 63.44% | 30.22% | 34.96% | |
| T0031 | 29.07% | 28.47% | 30.85% | 29.34% | 89.42% | 47.01% | 22.98% | 94.63% | 29.36% | 56.40% | |
| T0032 | 29.07% | 29.46% | 32.40% | 29.34% | 82.59% | 23.21% | 26.95% | 72.50% | 27.32% | 26.67% | |
| T0036 | 29.07% | 28.47% | 30.85% | 29.34% | 66.03% | 20.82% | 22.98% | 60.89% | 17.11% | 23.18% | |
| T0048 | 29.07% | 28.47% | 30.85% | 29.34% | 56.20% | 32.37% | 22.98% | 75.20% | 17.11% | 23.18% | |
| T0057 | 29.07% | 28.47% | 30.85% | 29.34% | 92.12% | 20.82% | 28.62% | 78.34% | 19.60% | 23.18% | |
| T0098 | 29.07% | 28.47% | 33.84% | 29.34% | 77.20% | 22.05% | 36.72% | 74.74% | 27.83% | 26.83% | |
| T0112 | 29.07% | 28.47% | 39.69% | 29.34% | 62.24% | 60.91% | 37.12% | 77.30% | 49.02% | 50.09% | |
| T0150A | 30.03% | 29.64% | 33.33% | 30.98% | 57.49% | 35.65% | 32.35% | 59.69% | 44.95% | 32.87% | |
| T0150B | 29.07% | 28.47% | 30.85% | 29.34% | 90.58% | 20.82% | 22.98% | 79.70% | 34.60% | 23.18% | |
| T0603 | 29.69% | 28.47% | 33.26% | 29.34% | 75.66% | 28.88% | 30.77% | 76.69% | 25.87% | 35.11% | |
| T0607 | 29.07% | 28.47% | 36.63% | 40.01% | 40.94% | 34.02% | 26.57% | 46.82% | 38.03% | 29.76% | |
| Lokakṣema | T0224 | 46.09% | 57.73% | 51.25% | 59.22% | 40.15% | 64.21% | 93.69% | 43.25% | 56.12% | 42.04% |
| T0280 | 31.94% | 94.08% | 98.56% | 45.29% | 27.38% | 97.86% | 39.66% | 32.28% | 90.10% | 54.77% | |
| T0417 | 56.05% | 73.50% | 51.40% | 76.73% | 44.03% | 65.97% | 56.31% | 53.19% | 66.19% | 73.55% | |
| Zhi Qian | T0054 | 61.33% | 57.31% | 87.32% | 29.34% | 31.44% | 49.85% | 51.59% | 54.16% | 60.63% | 81.60% |
| T0068 | 61.66% | 52.03% | 68.49% | 55.16% | 35.25% | 54.11% | 63.11% | 38.49% | 57.14% | 52.36% | |
| T0076 | 92.93% | 70.20% | 57.07% | 29.34% | 30.63% | 36.31% | 78.03% | 29.91% | 84.56% | 57.84% | |
| T0185 | 63.69% | 65.90% | 60.22% | 64.56% | 36.08% | 56.95% | 67.89% | 38.31% | 62.58% | 60.64% | |
| T0281 | 56.21% | 38.91% | 62.61% | 58.18% | 24.32% | 51.56% | 58.79% | 16.64% | 63.59% | 41.83% | |
| T0474 | 62.34% | 69.91% | 59.78% | 60.86% | 39.22% | 69.31% | 64.37% | 41.67% | 58.66% | 68.66% | |
| Kang Senghui | T0152 | 87.55% | 68.88% | 42.92% | 41.92% | 31.83% | 53.82% | 86.78% | 28.18% | 58.39% | 73.03% |
| Kang Mengxiang | T0137 | 36.04% | 56.41% | 37.47% | 82.46% | 25.33% | 65.30% | 48.04% | 15.28% | 90.45% | 94.17% |
| T0197 | 47.94% | 45.09% | 65.52% | 87.66% | 29.00% | 57.32% | 41.18% | 30.61% | 67.66% | 74.63% | |
| Dharmarakṣa | T0154 | 77.08% | 73.08% | 53.53% | 80.07% | 35.28% | 66.23% | 60.52% | 33.17% | 54.90% | 80.14% |
| T0186 | 70.49% | 77.27% | 73.62% | 86.61% | 29.95% | 67.92% | 51.31% | 28.54% | 59.57% | 74.02% | |
| T0222 | 56.91% | 57.02% | 43.77% | 75.18% | 45.74% | 72.51% | 90.70% | 41.77% | 55.97% | 62.38% | |
| T0263 | 80.18% | 78.50% | 71.62% | 89.11% | 30.68% | 69.20% | 58.98% | 28.58% | 72.41% | 74.82% | |
| T0398 | 78.41% | 89.28% | 53.07% | 77.98% | 35.01% | 70.26% | 57.20% | 29.93% | 77.29% | 66.11% |
Table C2. Aggregated stylometric score distribution across each translator.
| Author | func_之 | func_其 | func_皆 | func_及 | pos2_v_d | pos2_n_r | pos3_v_u_n | pos3_d_v_d | pos3_n_u_v | pos4_n_v_r_n |
| Yan Fotiao | 41.10% | 28.47% | 46.27% | 29.34% | 37.14% | 31.10% | 50.81% | 39.18% | 24.77% | 23.18% |
| An Shigao | 29.19% | 28.64% | 32.80% | 30.42% | 70.70% | 31.71% | 28.43% | 71.88% | 29.58% | 31.73% |
| Kang Senghui | 87.55% | 68.88% | 42.92% | 41.92% | 31.83% | 53.82% | 86.78% | 28.18% | 58.39% | 73.03% |
| Kang Mengxiang | 41.99% | 50.75% | 51.49% | 85.06% | 27.17% | 61.31% | 44.61% | 22.94% | 79.05% | 84.40% |
| Lokakṣema | 44.70% | 75.10% | 67.07% | 60.41% | 37.19% | 76.01% | 63.22% | 42.91% | 70.80% | 56.79% |
| Zhi Qian | 66.36% | 59.04% | 65.91% | 49.57% | 32.82% | 53.02% | 63.96% | 36.53% | 64.53% | 60.49% |
| Dharmarakṣa | 72.61% | 75.03% | 59.12% | 81.79% | 35.33% | 69.22% | 63.74% | 32.40% | 64.03% | 71.49% |
Table C3. Classification of An Shigao’s disputed translations.
| CBETA ID | Our Result | BERT Model Result (Lu et al. 2025) |
| T0101 | Yes | Yes |
| T0105 | No | No |
| T0109 | No | No |
| T0604 | No | No |
| T0605 | No | No |
| T0792 | No | No |
| T1557 | An Shigao | An Shigao |
| T0397_059 | No | No |