Daejeon, July 27–31
For years, digital humanities has been immersed in a research paradigm dominated by data mining and pattern recognition. While highly effective in revealing macro-level trend and correlations, this paradigm carries an inherent risk of “flattening” complex cultural meanings. In cultural heritage research, this risk is particularly acute: the digitization process often transforms rich, polysemic, and context-dependent heritage into standardized, structured data suitable for computational processing, inevitably simplifying the original ambiguities, contradictions, and situational specificities of culture. Consequently, subsequent visualizations or digital narratives, though visually compelling, often lack depth and thickness in cultural interpretation. Existing research in digital cultural heritage predominantly focuses on the “starting point” of digitization and the “endpoint” of presentation, while overlooking the crucial intermediate stage of deep interpretation of cultural heritage’s intrinsic meanings.
This study directly addresses this fundamental challenge. We argue for methodological innovation that leverages small-scale yet high-density, high-quality datasets to achieve deep and critical cultural interpretation. To this end, we systematically reconceptualize one of the core methodologies “thick description” from anthropology in the digital humanities context as “digital thick description.” Through a digital thick description of specific heritage objects, this research transcends surface-level pattern extraction to penetrate webs of meaning, effectively compensating the cultural oversimplification and bias inherent in the big data paradigm. It thus offers a new research path to digital humanities that emphasizes interpretive depth, contextual restoration, and cultural subjectivity. This is not only an intrinsic requirement for deepening cultural heritage studies, but also a theoretical response to the development of digital humanities, particularly providing a feasible, depth-focused (rather than scale-driven) paradigm for resource-scarce or highly specialized cultural domains, such as marginal or local heritage studies.
Existing research studies this issue mainly from three dimensions. The first dimension is digital humanities research on cultural heritage. Its evolutionary trajectory is clear: from foundational resource digitization and database construction (Zhou et al., 2017), to knowledge organization via metadata, ontologies, and knowledge graphs (Yang et al., 2024), and finally to visual presentation and service innovation using GIS (Gkadolou / Prastacos, 2021), 3D modeling (Wang et al., 2022), and digital storytelling (Yu, 2024). This lineage demonstrates the powerful capacity of digital humanities for data integration and processing, however, it largely sidesteps critical heritage studies’ emphasis on power asymmetries, colonial legacies, and contested meanings (Smith, 2006), as well as postcolonial digital heritage scholarship’s attention to marginalized voices and decolonial praxis (Risam, 2019). Consequently, it lacks dedicated methodological attention to how cultural meanings embedded in data—particularly small-scale data—are understood, translated, and interpreted. The second dimension is the application of thick description theory in cultural heritage studies. Since Clifford Geertz’s formulation, thick description has emphasized contextualized analysis to penetrate cultural surfaces and reveal the webs of meaning behind actions. In cultural heritage, scholars have applied it to interpreting intangible heritage practices (Kole, 2010) or material remains (Quick, 2020), uncovering underlying power relations, identity politics, and social structures. However, these studies largely follow traditional anthropological paradigms, relying on physical fieldwork and qualitative textual analysis, showing significant limitations when confronting already-digitized, multi-sourced, and heterogeneous heritage resources. The third dimension is discussions on thick description in the digital and AI era. Serra (2017) applied “thick description” to architectural heritage mapping, revealing the cognitive value of measured drawings as interpretive cultural texts. Wigard and Singh (2019) used “digital thick description” to unpack cultural plenitude in Mughal visual and textual sources. Tóth-Czifra (2019) noted that the complexity of cultural heritage data and the simplifying tendency of standardized metadata strip away cultural contexts. Pretnar Žagar and Podjed (2024) discussed how thick description and big data computational analysis can dialogue and transform in anthropological research. Kommers et al. (2025) proposed that LLMs can automatically generate thick description texts, offering ideas for digital realization. While recognizing thick description’s importance in the digital age, these studies have yet to systematically construct a methodological framework that deeply integrates digital humanities’ computational methods with thick description’s interpretive logic. Consequently, a critical research gap emerges: how to systematically transplant and reconstruct the classic deep-interpretation methodology of thick description within digital humanities’ technological environments and data ecosystems, enabling it to handle the complexity of digitized heritage while preserving its theoretical core that pursues meanings in depth, thereby providing methodological possibilities for small data’s deep interpretation in digital humanities.
To address these issues, this study first systematically reviews Geertz’s thick description theory (Geertz, 1973), demonstrating its viability and necessity for cultural heritage research, and summarizes traditional thick description stages. Building upon these, we propose and systematically construct a “Digital Thick Description of Cultural Heritage” framework, integrating digital humanities’ basic research types (McCarty, 2001). Its core lies in translating Geertz’s two key features of thick description—contextualization and stratification—into operational digital humanities workflows. The framework comprises four iterative core stages (Figure 1):
(1) Digital Field Immersion: Transcending the physical constraints of traditional fieldwork, this stage aggregates multi-sourced, multimodal data around target heritage to construct a “digital field.” Data encompass foundational substrate (geometric, textural, and structural information of heritage), historical substrate (relevant archives, local gazetteers, historical maps), and social substrate (newspapers, books, academic works). After digitization and standardized integration, a computable and immersive digital field is formed. Researchers observe through “digital presence,” producing “digital fieldnotes.”
(2) Cultural Node Selection: From digital fieldnotes, nodes with thick description potential are identified as basic analytical units. Drawing from the CIDOC-CRM model and related international heritage documents, six cultural elements are induced: material, temporal, spatial, actor, event, and symbol. These elements are extracted from digital fieldnotes. Through a combination of manual close reading (as the primary screening mechanism) and LLM-assisted analysis, “cultural nodes” are selected based on thick description criteria such as cultural salience and meaning concentration.
(3) Contextual Layered Interpretation: This is the core of digital thick description. Geertz’s thick description theory is transformed into an operational three-layer topology: the Surface Layer describes heritage’s superficial phenomena, the Contextual Layer constructs associated contexts, and the Meaning Layer reaches deep cultural logics. Starting from a cultural node in the Surface Layer, appropriate nodes in the Contextual Layer are selected and connected as context; completing a connection to the Meaning Layer finalizes a node’s interpretation. This process makes the implicit interpretive layers of traditional thick description explicit as concrete topological paths, enhancing the visibility and operability of cultural meaning interpretation. By planning and executing different interpretive paths, structured “interpretive prototypes” are generated.
(4) Visual Ethnography Generation: Interpretive prototypes are transformed into “visual digital ethnographies of cultural heritage” that integrate visualization charts with thick description texts. Using visualization techniques such as GIS and social network analysis to present “interpretive prototypes,” LLMs are introduced as auxiliary drafting tools rather than autonomous interpreters: expert scholars retain ultimate authority to critically evaluate, refine, and validate generated texts. Through prompts designed based on the “MIRACLE” thick description evaluation framework (Younas et al., 2023), LLMs generate preliminary thick description texts constrained by visualization charts, scholars then scrutinize these drafts to mitigate risks of algorithmic bias, ultimately synthesizing expert-validated ethnographies.
Figure 1. Digital thick description of cultural heritage framework.
To validate this framework, this study conducted a digital thick description practice on the Bibliotheca Zikawei (徐家汇藏书楼) in Shanghai, a heritage building founded by French Jesuit missionaries in 1847. As one of China’s earliest modern libraries, it witnessed modern Sino-Western cultural exchanges. Integrating data from Shen Bao (申报), relevant local gazetteers, and OCLC catalogs, we constructed a digital field and selected three cultural nodes—“Geographical Location of Bibliotheca Zikawei”, “Catholic Literature Collection of Bibliotheca Zikawei” and “Bibliotheca Zikawei’s Foreign-Language Name”—for multi-layered contextual interpretation, generating a visual ethnography integrating maps, network graphs, timelines, and thick description texts (Figure 2a).
Figure 2. Digital thick description on the Bibliotheca Zikawei. (a) Digital thick description path for Bibliotheca Zikawei. (b) Visualization of digital thick description results for “Geographical Location of Bibliotheca Zikawei” node. (c) Visualization of digital thick description results for “Catholic Literature Collection” node. (d) Visualization of digital thick description results for “Bibliotheca Zikawei’s Foreign-Language Name” node.
The thick description of “Geographical Location of Bibliotheca Zikawei” reveals that its site at the edge of the old Shanghai concession, correlated with density distributions of over 40 libraries in Shanghai and early 20th-century population density maps (Figure 2b), exposes the Jesuits’ cultural strategy: establishing an enclave of knowledge production within the interstices of power. The “Catholic Literature Collection” node, through author collaboration networks and national-origin analysis, shows the collection was not Western unilateral imposition but rather “co-shaped” by Chinese literati like Xu Guangqi (徐光启) and missionaries like Matteo Ricci (Figure 2c), demonstrating adaptive reconstruction of Catholic knowledge within the Chinese context. The “Foreign-Language Name” node traces multilingual name changes across 1,005 temporal bibliographic data slices (Figure 2d), revealing the shift from “Zikawei” (French/Latin-dominated) to “Xujiahui” (rise of Pinyin)—where name selection becomes a symbolic arena for identity politics and knowledge power struggles, mapping the micro-processes of transition from colonial discourse to modern national identity. The three nodes collectively demonstrate how Bibliotheca Zikawei persistently functioned as a “cultural intermediary” in China’s knowledge system transformation from late Qing to modern times, weaving webs of meaning amid tensions between local and Western, religious and secular forces.
This digital thick description practice on the Bibliotheca Zikawei demonstrates that through deep interpretation of small-scale yet high-density datasets, we can reveal nuanced yet crucial cultural meanings easily obscured in large-scale data analysis. These interpretive achievements surpass mere factual statements or pattern descriptions, entering the realm of meaning construction and critical interpretation. The value of the Digital Thick Description framework lies not only in providing operational tools for cultural heritage studies but also in opening a possible paradigm for digital humanities beyond big data patterns highlighting small data’s unique advantages in countering algorithmic bias, safeguarding cultural differences, and empowering marginal research.
This study’s contributions are twofold: First, it offers a feasible path to combine humanities’ deep-interpretation tradition with digital technological capabilities, proving the immense potential and unique value of “small data” deep analysis in digital humanities. Second, by operationalizing thick description, it provides a systematic, meaning-generation-process-focused interpretive toolkit for cultural heritage studies and broader digital humanities disciplines.