Daejeon, July 27–31
Automatic detection of text reuse in texts written in varieties of what is known as Literary Chinese, Hanmun or Kanbun has become an important topic in the field of East Asian Studies and beyond, and a range of related tools serving different related purposes have been developed. Thus, Michael Radich’s TACL identifies intersections of character n-grams to establish usage profiles for individual texts that then can be employed for authorship profiling based on reuse of similar expressions or lack thereof (cf., e.g., Radich 2017). The Chinese Text Project and Dharma Nexus (Sturgeon 2018a; 2018b and Nehrdich 2020) identify matches between a given text and a given corpus based on n-gram shingling. Optimized for large corpora, these solutions cluster reoccurring textual passages that appear across multiple works, effectively allowing to trace philosophical memes circulating through a tradition. Most recently, the LLM based MITRA-zh/ Dharmamitra (Nehrdich a.o. 2023), even facilitates searches of semantically similar passages across languages. Finally, pairwise comparison methods such as those employed by Vierthaler and Meis (2019) and in our own MNGRAM software allow for fine-grained alignments and analysis of parallel passages between two given texts.
While thus a range of software solutions allowing for the identification of instances of text reuse in texts written in traditional Chinese does exist, the problem of identifying the direction of text reuse hitherto has been neglected. This is deplorable for two reasons: For one, quite a few traditional authorship ascriptions remain disputed, and (as assertive as impressionistic) claims of “strong influence” are, if at all, usually backed up by content level arguments that tend to remain inconclusive. For another, even philological research questioning traditional ascriptions of texts to this or that author remains informed by the naïve conception of solitary authorship. For instance, scholars have identified considerable parallels between the Kisillon so 起信論疏 (T.1844) and the Dasheng qixin lun yiji 大乘起信論義記 (T.1846), two commentaries traditionally ascribed to Wŏnhyo (617-686) and Fazang (643-712), resp. The direction of this obvious text reuse is usually tacitly taken for granted solely based on the purported solitary authors’ biographical dates, the possibility of a (partially) wrong ascription of a text that actually has been compiled or heavily edited by some disciple(s) usually not even being taken into consideration.
There have been attempts to tackle the problem of text reuse direction in other languages. Notably, Shrestha and Solorio (2015) compare bag of word n-gram profiles of shared passages with segments of the remaining text(s). The text with the greatest average overlap then is identified as the original one. Grozea and Popescu (2010) proceed from the observation that vocabulary of the shared passages should appear earlier in the original text, which in plots of matches between the files results in horizontal or vertical stripes with a higher density of matches. Thus, Grozea and Popescu compute the contrast between these regions and neighboring ones, the file with the larger contrast then being considered the original one.
In our own approach, we will grosso modo follow the approach taken by Shrestha and Solorio, at the same time, however, incorporating the basic insight of Grozea and Popescu: By means of our locally written MNGRAM aligner, we will identify substantial instances of textual reuse, i.e. alignments of a minimum length allowing for a meaningful bag of words comparison of the n-grams appearing in these sections with the rest of the text(s). The actual comparisons will employ CJK character n-grams of different lengths and also otherwise differ in two respects from Shresha and Solorio (2015): For one, taking not only Grozea and Popescu’s (2010) interesting observation into consideration, but also the directedness of the traditional writing process before the advent of cut & paste, the answer to the question whether a given section in a text organically “fits” the text it occurs in usually should depend primarily on the segments preceding the passage in question. Hence we weigh matches with n-grams in earlier segments higher than matches with n-grams in subsequent ones. For another, and more importantly, we have to consider camouflage through adaption of reoccurring stylistic patterns. Thus, in contrast to Shrestha and Solorio (2015), we do not concentrate on the “top” matches with the highest frequency, but include less often appearing matches and even prioritize these low frequency ones in our weighing formula.
Due to the non-availability of substantial corpora of texts with undisputed authorship, a validation comparable to the standards of analyses based on Western languages is out of reach. However, by evaluating the consistency of correct ascription of matching segments between texts with known textual relations, at least some conclusions on the performance of the method may be drawn. Thus, we will evaluate the performance of our method based on the relation of the matching and non-matching parts of the two recensions of the Dasheng qixin lun 大乘起信論 (T.1666 and T.1667) attributed to Paramārtha (Zhendi 真諦, 499-569) and Śikṣānanda (Shichanantuo 實叉難陀, 652-710), resp., as well as that of the Dasheng qixin lun yishu 大乘起信論義疏, traditionally attributed to Huiyuan 慧遠 (523~592), against the already mentioned Kisillon so 起信論疏 and Dasheng qixin lun yiji 大乘起信論義記. (Although the former text has been shown to contain notes that prove that it has been at least edited by Huiyuan’s 慧遠 disciples (Lai 1975), it evidentially antedates the latter texts.)
After verifying the reliability of our approach, we will then employ it to validate the direction of textual borrowing between Fazang’s Dasheng qixin lun yiji 大乘起信論義記 and Wŏnhyo’s Kisillon so 起信論疏.