Daejeon, July 27–31
1. Introduction
Language is among the most sensitive semiotic systems for reflecting social change. In China, the rapid growth of digital platforms such as Weibo, WeChat, and Douyin has accelerated the diffusion of new buzzwords, including ‘内卷’ (involution/hyper-competition), ‘躺平’ (lying flat/opting out of competition), and ‘打工人’ (self-deprecating worker identity). These buzzwords serve as linguistic indicators of what a society attends to, what emotions it shares, and how it interprets reality.
Large-scale corpora provide a foundation for empirically analyzing the emergence, diffusion, and decline of buzzwords, yet a fundamental methodological problem exists: high corpus frequency does not necessarily correspond to buzzword usage. For instance, ‘谷子’ was selected as a 2025 buzzword meaning 'goods' (a phonetic borrowing), but the overwhelming majority of its BCC corpus tokens refer to its established meaning of 'grain.' Conversely, slot-based constructions such as ‘多巴胺 XX’ may be structurally absent from standard corpora.
This study departs from this problem and poses three research questions: (RQ1) To what extent does raw corpus frequency reflect actual buzzword-meaning usage? (RQ2) Does M0/M1 disambiguation and frequency correction change lifecycle analysis outcomes for semantic-shift buzzwords? (RQ3) How do word formation type and corpus reliability relate to lifecycle type, and what are the key factors explaining lifecycle trajectories?
2. Research Design
2.1. Data and Materials
This study analyzes 110 "Annual Buzzwords" selected by the Chinese magazine 『咬文嚼字』(Yaowen Jiaozi) from 2015 to 2025. This list is compiled through an expert panel using consistent criteria each year, thereby avoiding researcher selection bias and covering 11 years of diachronic variation.
The primary analytical resource is the BCC Corpus (BLCU Corpus Center, approximately 15 billion characters), which includes newspapers, literature, spoken language, and other genres. While BCC does not represent the full spectrum of internet usage, internet-only buzzwords (e.g., 996, city-不-city) that are structurally absent from the corpus are treated as an analytical finding in themselves. The final analysis covers 95 keywords and 130,388 concordance tokens collected from BCC.
2.2. Analysis Pipeline
[Figure 1] Analysis Pipeline for Corpus Capture and Lifecycle of Chinese Buzzwords
The pipeline runs word formation analysis for all 110 keywords and M0/M1 semantic disambiguation for 29 semantic-shift keywords in parallel, then integrates results at the co-occurrence analysis and lifecycle typology stages. Key steps include: (1) search term processing and BCC token collection, (2) time-series reliability classification (Grades A/B/C/D), (3) M0/M1 disambiguation and corrected frequency calculation, (4) co-occurrence extraction and MI/LL-based clustering, (5) half-life and settlement index-based lifecycle typology (Flash/Sedimentation/Resurrection), and (6) chi-square statistical testing. All analyses were implemented in Python-based Colab notebooks.
3. Semantic Reliability Verification of Corpus Frequency
3.1. Reliability Grade Classification
Buzzword tokens retrieved from BCC did not exhibit uniform reliability. Keywords were classified into five grades based on time-series frequency patterns: A (high reliability: M1 meaning dominant), B (mixed: M0 and M1 coexist), C (low reliability: M0 dominant), D (low capture: structurally underrepresented), and D_nodata (insufficient data).
[Figure 2] BCC Corpus Reliability Grade Distribution (112 keywords)
(There are 110 buzzwords selected for 『咬文嚼字』 but the analysis unit is organized into 112 keywords during the actual corpus search and variant separation process.)
While Grade A accounts for the largest share (44 keywords), the remaining B, C, D, and D_nodata grades total 68 keywords. This demonstrates that BCC raw frequency is a useful starting point but cannot be directly interpreted as buzzword frequency without semantic verification.
3.2. M0/M1 Disambiguation and Corrected Frequency
For 29 semantic-shift keywords, a total of 648 concordance tokens were manually reviewed. The results showed: original meaning (M0) 390 tokens (60.2%), buzzword meaning (M1) 245 tokens (37.8%), and excluded (N) 13 tokens (2.0%). This confirms that established-meaning tokens constitute over half of the raw frequency for semantic-shift buzzwords.
[Figure 3] Raw Frequency (gray) vs. Corrected Frequency (red) for the 6 Most Contaminated Keywords
The figure shows cases such as 神兽 (M1 ratio: 3.0%), 尬 (4.5%), 锦鲤 (11.5%), 拿捏 (12.9%), 天花板 (13.3%), and 大白 (13.6%). Using raw frequency without correction would severely overestimate actual buzzword usage, confirming that semantic disambiguation and frequency correction are essential in corpus-based buzzword research.
3.3. Co-occurrence-Based Semantic Field Analysis
Co-occurrence analysis revealed five semantic clusters among the buzzwords: (1) policy and institutional discourse (中国式现代化, 新质生产力, etc.), (2) youth sentiment and labor discourse (内卷, 躺平, 打工人, etc.), (3) platform and internet culture (流量, 破防, 尬, etc.), (4) technology and digital discourse (人工智能大模型, 数智化, etc.), and (5) everyday life and relational expressions (宝宝, 搭子, etc.). This demonstrates that buzzword meaning is shaped not by dictionary definitions alone, but by the co-occurring words and socio-generic contexts in which they are used.
4. Lifecycle Analysis Using Half-Life Concept
4.1. Lifecycle Type Classification
To analyze temporal change, buzzwords were classified into four lifecycle types using half-life concept: Flash (rapid rise followed by decline), Sedimentation (sustained presence), Resurrection (reactivation after decline), and Too_Early (insufficient observation period). The half-life is defined as the time from peak frequency to the point at which frequency drops below half the peak value.
[Figure 4] Representative Cases by Lifecycle Type
Flash-type buzzwords are prominent among event- or meme-driven expressions (供给侧, 不忘初心, etc.). Sedimentation-type buzzwords appear in expressions linked to policy or social-issue discourse (获得感, 命运共同体, etc.). Resurrection-type buzzwords are reactivated through re-contextualization (新质生产力, 中国式现代化, etc.). Too_Early keywords are those selected in 2024–2025 with insufficient observation periods.
4.2. Lifecycle Comparison: Raw vs. Corrected Frequency
[Figure 5] Lifecycle Type Change: Raw Frequency (top) vs. M1 Corrected (bottom)
Semantic correction changed the lifecycle type for some buzzwords. Tianhuaban shifted from Resurrection (raw) to Flash (corrected). Dabai likewise changed from Sedimentation (raw) to Flash (corrected), as its pandemic-era M1 usage was concentrated in a specific period. In contrast, dui maintained its Sedimentation type even after correction, showing that contamination rate alone does not uniformly determine lifecycle type.
4.3. Word Formation and Lifecycle
[Figure 6] Keyword Count (left) and BCC Frequency Distribution (right) by Word Formation Type
Among the 110 buzzwords, compounds were the most frequent (33, 29%), followed by semantic-shift (28, 25%), phrase/sentence (24, 21%), and derivation (14, 12%). Compounds and semantic-shift types showed relatively higher BCC frequencies, while phrase/sentence and derivation types exhibited greater item-level variance. However, neither the word formation × lifecycle type chi-square test (χ² = 20.76, p = 0.2919) nor the reliability grade × lifecycle type test (χ² = 1.57, p = 0.6668) was statistically significant. This suggests that buzzword lifecycle is more strongly explained by discursive reproducibility than by word formation type itself.
5. Discussion and Conclusion
This study examined how the BCC Corpus captures 110 Chinese buzzwords selected by 『咬文嚼字』 (Yaowen Jiaozi). The key findings are as follows.
First, corpus frequency ≠ buzzword frequency. Excluding the 44 Grade-A keywords, 68 keywords showed limitations in directly interpreting raw frequency. Among the 648 tokens of 29 semantic-shift keywords, M0 (original meaning) accounted for 60.2%, indicating that raw frequency-based analysis can overestimate buzzword usage.
Second, semantic correction changes lifecycle types. As confirmed in the cases of 天花板 (Resurrection → Flash) and 大白 (Sedimentation → Flash), raw and corrected frequencies can yield different lifecycle classifications, demonstrating that semantic correction is essential for lifecycle analysis.
Third, the key factor governing buzzword lifecycle is discursive reproducibility, not word formation. Event- and meme-driven expressions tend toward Flash-type decline, while expressions linked to policy or social-issue discourse tend toward Sedimentation. Neither the word formation × lifecycle nor reliability grade × lifecycle tests were statistically significant.
The methodological contribution of this study lies in treating corpus frequency not as a given result but as data requiring verification. The semantic verification pipeline—combining reliability grading, M0/M1 disambiguation, corrected frequency calculation, co-occurrence analysis, and lifecycle typology—can be extended as a digital humanities methodology applicable not only to Chinese buzzwords but also to neologisms, meme expressions, and social keyword analysis more broadly.
Limitations include BCC's inability to represent the full range of internet platform usage, the irreducibility of human judgment in semantic disambiguation, and small sample sizes in some word formation categories. Future research should expand the data scope to include social media, news, search indices, and multiple corpora, and should focus on analyzing the syntactic function shifts that occur as semantic-shift buzzwords transition from M0 to M1 meanings.
Keywords: Digital Humanities, Chinese Buzzwords, BCC Corpus, Semantic Reliability, Lexical Lifecycle, Word Formation