DH 2026

Daejeon, July 27–31

Wed, July 2916:30–18:00S042206-208
Long Paper

The Construction of Records of the Grand Historian Digital Humanities Knowledge Base: Engaging Cultural Memory, Linguistic knowledge and GIS Service

Yue Zhu
School of Chinese Language and Literature, Nanjing Normal University, No.122, Ninghai Road, Nanjing 210097, China · 2771130171@qq.com
Tongzheheng Zheng
Shangrao Key Project Service Center, No. 2 Jinxiu Road, Shangrao 334000, China; Center of Language Big Data and Computational Humanities, Nanjing Normal University, No.122, Ninghai Road, Nanjing 210097, China · 250583749@qq.com
Bin Li
School of Chinese Language and Literature, Nanjing Normal University, No.122, Ninghai Road, Nanjing 210097, China; Center of Language Big Data and Computational Humanities, Nanjing Normal University, No.122, Ninghai Road, Nanjing 210097, China · libin.njnu@gmail.com
Minxuan Feng
School of Chinese Language and Literature, Nanjing Normal University, No.122, Ninghai Road, Nanjing 210097, China; Center of Language Big Data and Computational Humanities, Nanjing Normal University, No.122, Ninghai Road, Nanjing 210097, China · fengminxuan@njnu.edu.cn

The Records of the Grand Historian (Shiji), complied by Sima Qian during the Western Han Dynasty, is a foundational monument of Chinese historiography and a work of enduring global significance: it is the first of “Twenty-four Histories(二十四史) and the earliest biographical general history in China, chronicling over three thousand years from the Yellow Emperor(traditionally dated to c. 2700 BCE) to Emperor Wu of the Western Han Dynasty(122 BC)(Chi Changhai 2021). Shiji integrates history, philosophy, literature, and ethics, providing important insights into ancient social structures, cultural identities, and humanistic thought.

However, its classical Chinese syntax, vast scale (130 chapters, over 520,000 characters), and dense contextual complexity have long hindered global accessibility. Moreover, traditional research on shiji has relied primarily on linear, text-by-text reading practices, which hinders large-scale text analysis and the construction of structured and dynamic textual correlations..

In response to these constraints, this study presents the Shiji Digital Humanities Knowledge Base (Shiji-DHBase), a structured, multi-layered knowledge infrastructure. Building on our prior achievements of the Basic Annals (本纪) (Li Bin et al 2020:4(2): 528-536)and Biographies(列传)(Zheng Tongzheheng et al 2022: 8(6), 40–55), this study integrates the above resources, supplements the Hereditary Houses (世家), and develops a systematic full-scale knowledge base for the complete Shiji. Drawing on advances in Classical Chinese information processing(Chen Xiaohe et al 2013)(Fang Wenjie 2023) and digital humanities(Schreibman S et al 2004), this platform extracts valid information from ancient texts, establishes structured and dynamic connections between contents, expands new perspectives for the study of vocabulary, figures and locations in the Shiji, and provides a new method for constructing digital knowledge bases of biographical history books.

Methodologically, the project integrates Nanjing Normal University’s Language Big Data and Computational Humanities Center expertise in corpus construction(Shimin et al 2010:24(02): 39-45), semantic annotation, and low-resource language technology with cutting-edge digital humanities tools. First, we adopted the punctuated and revised Shiji published by Zhonghua Book Company for unified data sorting, ensuring the authenticity and standardization of the original text(Shiji Revision Group 2013). Based on this textual foundation , We applied automated tools for Classical Chinese word segmentation and part-of-speech tagging, followed by manual correction to ensure high-quality linguistic annotation. To further transform the text into structured data suitable for computational analysis, we performed entity annotation focusing on core nouns, including historical figures and places. The overall annotation workflow is shown in Table 1.

Table 1 data annotation workflow

Annotation workflowExamples of annotation
Original text

項梁使沛公及項羽別攻城陽,屠之。

(Xiang Liang ordered Pei Gong and Xiang Yu to separately attack Chengyang, and they captured and sacked the city.)

Automatic word segment and part-of-speech tagging 項梁/nr 使/v 沛公/nr 及/c 項羽/nr 別/w 攻/v 城陽/ns ,/w 屠/v 之/r 。/w
Manual verification項梁/nr 使/v 沛公/nr 及/c 項羽/nr 別/d 攻/v 城陽/ns ,/w 屠/v 之/r 。/w
Manual entity annotation項梁/nr[Person ID:3076] 使/v 沛公/nr/[Person ID:2621] 及/c 項羽/nr[Person ID:3079] 別/d 攻/v 城陽/ns[Location ID:5114] ,/w 屠/v 之/r 。/w

For named entities, we assigned unique identifiers (IDs) to historical figures and defined a person-entity schema that includes person ID, standard name, alternative names, gender, and associated dynasty. Similarly, we established a parallel place-entity schema consisting of place ID, ancient name, modern equivalent, place category (e.g. administrative units, feudal states, rivers, mountains), and GIS coordinates (Baidu Maps, a widely used Chinese web mapping service), restoring the historical spatiotemporal background of the text. To identify modern equivalents and geographic coordinates, we consulted philological references such as Shiji Diming Kao(史记地名考) to verify the historical geography(Qian Mu 2001). Examples of person and place entities are shown in Tables 2 and 3.

Table 2 examples of entities of persons

Person IDCanonical NameAlternative NamesGenderHistorical Period
2621漢高祖(Emperor Gaozu of Han)漢王(King of Han), 高祖(Gaozu), 沛公(Pei Gong), 高帝(Gao Emperor), 高皇帝 (Gao Emperor), 劉季(Liu Ji), 季(Ji), 赤帝(Chi Emperor), 沛令(Magistrate of Pei), 武安侯(Marquis of Wu’an), 安武侯((Marquis of Anwu), 高皇( Gao Emperor), 漢高帝(Gao Emperor of Han), 漢太祖(Emperor Taizu of Han), 太上皇(Retired Emperor), 太祖(Emperor Taizu), 祖(The Founder) MaleLate Qin to Early Han

Table 3 examples of entities of places

Place ID Place NameModern EquivalentPlace CategoryGIS Coordinates
5114城陽(Chengyang)Qingdao, Shandong Province, ChinaQin-era toponym120.387145,36.072825

After performing word segmentation, part-of-speech tagging, and entity information annotation on Shiji, we obtained a text table, a text annotation table, a person table, and a place table. By linking the person table and the place table and exploiting co-occurrence information, we derived both person–person and person–place co-occurrence tables. These six one-dimensional sequence tables are then integrated into a multidimensional knowledge framework, organizing texts, vocabularies, and entities in a unified structure. Based on this framework, we develop a structured and dynamic digital humanities knowledge base for Shiji. The overall architecture is illustrated in Figure 1.

Figure 1. Architecture of the Shiji Digital Humanities Knowledge Base

Building on Shiji-DHBase, we conducted quantitative analyses and visualizations of Shiji at multiple levels, including vocabulary, character entities, and place entities. It shows that the Shiji text is dominated by monosyllabic words (85.09%), with a small proportion of multi-character items, reflecting typical Old Chinese features and the need for segmentation. In addition, the quantitative analysis identifies 3,445 character entities(with the Hereditary Houses being the most character-dense section) and 1,954 place entities(most of which appear in the Biographies section), along with their frequencies of appearance and inter-entity relations. Based on these data, we have developed visualizations that reveal a core network of key figures on Liu Bang(刘邦)--connecting him to Xiang Yu(项羽), Zhang Liang(张良), Han Xin(韩信), and others. High-frequency character-location co-occurrences(e.g. Liu Bang- Yingyang, Xiang Yu – Pengcheng) further support a narrative structure : core landmarks such as the Yellow River and Yingyang, together with the movement of Liu Bang’s group, collectively shape the main storyline of Shiji.

This study reinterprets Shiji—a foundational work of Chinese culture—not as a static archival artifact but as a dynamic corpus for research and public engagement. To support this, we have developed a web platform offering full-text, pattern, person, and place search. Pattern queries allow users to search part-of-speech information or collocations associated with specific grammatical categories, while person search allows users to trace an individual directly in the text or examine their relationships. To highlight the spatial narratives in Shiji, Shiji‑DHBase integrates a Baidu Maps–powered visualization tool for exploring the text’s historical geography. For example, when users enter a historical figure’s name, the platform will automatically generate the person’s movement trajectory across ancient states and aligns it with corresponding modern territorial boundaries. In addition, we developed a dedicated reading platform supporting word segmentation, traditional–simplified conversion, and lexical parsing where users can access biographical information, trace movements, and follow narrative distributions across the corpus.

Accordingly, we provide a user-oriented interface for different audiences. Linguists can explore classical function words and related phenomena, while historians can examine diplomatic and political relations through person co-occurrence networks. The platform also supports teaching, for instance by enabling interactive exercises such as mapping Liu Bang’s narrative geography, and offers map-based access to curated multimedia content for the general public. In addition, students are supported by a flexible reading environment with adjustable linguistic aids that help them overcome barriers to reading classical Chinese.

In all, by linking word segmentation, entity annotation, spatiotemporal reconstruction, and a retrieval system, Shiji-DHBase demonstrates how digital humanities can transform ancient texts into shared cultural resources rather than static relics. It also offers a case for reflecting on for whom cultural memory is preserved, and how technological tools, together with humanistic concerns, may reshape preservation toward broader public access. Going forward, we aim to integrate large language model–based translation to improve cross-linguistic accessibility and strengthen international communication. By expanding the knowledge base and encouraging wider participation, the platform further supports reflection on the use and significance of cultural heritage.

References
  1. Chen, Xiaohe / Feng, Minxuan / Xu, Runhua (2013): Information Processing for Pre-Qin Texts. Beijing: World Book Publishing Company.
  2. Chi, Changhai (2021): A Lexical Study of the Shiji. Hangzhou: Zhejiang University Press.
  3. Fan, Wenjie / Wang, Dongbo / Huang, Shuiqing (2023): Automatic sentence segmentation for classical Chinese: The Spring and Autumn Annals as an example, in: Digital Scholarship in the Humanities 38, 3: 1067–1077.
  4. Li, Bin / Li, Yaxin / Qian, Yang et al. (2020): From history book to digital humanities database: the Basic annals of the Shiji, in: Journal of Chinese History 4, 2: 528–536.
  5. Qian, Mu (2001): A Study of Place Names in the Shiji. Beijing: Commercial Press.
  6. Schreibman, S / Siemens, R / Unsworth, J (2004): A companion to digital humanities. Malden, MA: Blackwell Publishing Ltd.
  7. Shi, Min / Li, Bin / Chen, Xiaohe (2010): An Integrated CRF-Based Approach to Word Segmentation and Tagging for Pre-Qin Chinese, in: Journal of Chinese Information Processing 24, 2: 39–45.
  8. Shiji Revision Group (2013): Shiji: Collated and Revised Edition. Beijing: Zhonghua Book Company.
  9. Zheng, Tongzheheng / Li, Bin / Feng, Minxuan et al. (2022): Structured exploration of historical classics: Construction and visualization of a digital humanities knowledge base for the Shiji “Biographies”, in: Big Data 8, 6: 40–55.