DH 2026

Daejeon, July 27–31

Long Paper

A Methodology for YouTube Comments Topic Modeling: Limitations of Topic Reduction and Considerations for Human InterpretationA Cross-Linguistic Analysis of Meditation Discourse

Jihyun Lee
Graduate School of Culture Technology, KAIST, Daejeon, Republic of Korea · dwg07072@kaist.ac.kr
Jieun Woo
Graduate School of Culture Technology, KAIST, Daejeon, Republic of Korea · jieunwoo@kaist.ac.kr
Bong Gwan Jun
School of Digital Humanities and Computational Social Sciences, KAIST, Daejeon, Republic of Korea · junbg@kaist.ac.kr

Introduction

Scholars have long classified narratives and discourses by grouping similar plots, motifs, and themes. The Aarne–Thompson–Uther Index is a representative example of this tradition, classifying folktales by recurrent tale types and motifs (Uther 2004). Yet such work relied on expert manual comparison, which limited sample size and left fewer opportunities for systematic quantitative validation. Today, computational methods allow much larger-scale textual exploration.

In Digital Humanities, topic modeling has become a method for examining large-scale discourse (Blei et al. 2003; Grootendorst 2022; Meeks / Weingart 2012). This approach is particularly relevant to social media data, where large volumes of user-generated texts make manual reading difficult. Social media comments record contemporary emotions, participation, and cultural response (Thelwall 2018), but they are also heterogeneous and context-dependent, requiring interpretation rather than mechanical analysis (Grimmer / Stewart 2013; Marwick / Boyd 2011; Hine 2000). Topic modeling is thus exploratory, not a substitute for interpretation.

Previous studies have compared LDA and Top2Vec to evaluate topic model performance (Egger / Yu 2022). Recently, BERTopic has become prominent for producing context-aware clusters in short-text corpora (Kaur / Wallace 2024; Wang et al. 2024), leading to pipeline-optimization studies (Azher et al. 2024; Borčin / Jose 2024). However, statistical metrics used to evaluate topic models, such as coherence and diversity, do not necessarily make topics interpretable (Chang et al. 2009; Röder et al. 2015). Scholars have also noted that automated content analysis should be validated on each new dataset, since its performance can vary by context (Grimmer / Stewart 2013). This issue is especially relevant to social media texts, which need to be read in relation to the social, cultural, and communicative conditions in which they are produced (Roberts et al. 2014: 1068; Kozinets 2002: 65).

Content analysis emphasizes manifest and latent meanings (Berelson 1952), while thematic analysis identifies, reviews, and defines themes in relation to research questions (Braun / Clarke 2006). Unlike automated clustering, human interpretation can account for context, implication, symbolism, and subtle semantic differences. It can also reorganize model-generated topics into themes suited to the analytical purpose. Thus, human thematic interpretation is needed not only to validate topic models, but also to turn exploratory outputs into humanities research.

Topic and model selection have been examined in traditional models (Griffiths / Steyvers 2004; Taddy 2012), yet less is known about BERTopic in multilingual and multicultural datasets. It remains unclear how cross-linguistic preprocessing choices alter topics, how they should be decided, and when multilingual topics are comparable. This study addresses this gap by analyzing YouTube comments from Korea and the United States through BERTopic and human interpretation. It quantitatively demonstrates the usefulness of human thematic interpretation for data, and proposes a workflow prioritizing interpretability, transparency, and contextual validity over statistical stability alone.

Data and Methods

Focusing on the cultural nuances of South Korea and the United States meditation motivations, we analyzed comments from guided meditation videos on YouTube. To ensure cross-cultural comparability, the research design prioritizes a balanced framework that aligns disparate linguistic and communicative patterns into a singular interpretive structure. In accordance with the ethical guidelines of the Association of Internet Researchers (AoIR 2020), all personal identifiable information was anonymized, and quoted comments were modified and paraphrased to protect user privacy. A total of 56,767 comments were collected from U.S.-based YouTube channels and 41,279 comments from South Korean channels. To prevent the overrepresentation of highly active users, we limited the data to one comment per user. For users with multiple entries, the comment with the highest like count was retained; if like counts were identical, the longer comment was selected to maximize information density.

Text preprocessing is a critical factor in unsupervised text analysis (Denny / Spirling 2018; Schofield et al. 2017). After removing advertisements, spam, URLs, emojis, and mentions, we normalized whitespace and repeated characters. We then labeled comments as short, medium, long, or extremely long according to each language’s length distribution. Short reactive comments and extremely long outliers—up to 8k characters—were excluded, since short social media texts often lack sufficient co-occurrence and contextual information (Mehrotra et al. 2013; Churchill / Singh 2021). Medium-length comments were retained with minimal intervention, while long comments were segmented into sentences to improve thematic extraction.

We then modeled topics with BERTopic, which identifies thematic structures through transformer-based contextual embeddings, dimensionality reduction, and clustering. Contextualized document representations produce more coherent topics than traditional bag-of-words approaches (Bianchi et al. 2021). We used a multi-stage refinement procedure: initial topic extraction, removal of irrelevant or noise topics, and thematic aggregation (Cheng et al. 2022). For cross-cultural comparison, multilingual embeddings served as the main strategy; to check possible English-centric bias in multilingual models (Papadimitriou et al. 2023), each language was also modeled independently with language-specific embeddings. In HDBSCAN clustering, min_cluster_size was adjusted by document volume to avoid fragmented, uninterpretable micro-clusters (McInnes et al. 2017), and domain-specific stopword sets were matched across languages to ensure consistent topic representation.

We systematically compared five model configurations: 1) the primary-stage model, 2) the secondary-stage model, and three third-stage variants: 3) human-guided aggregation, 4) automated hierarchical aggregation, and 5) machine-led aggregation constrained to the number of thematic clusters identified through human-guided aggregation. Since topic granularity remains a central issue in topic modeling (Taddy 2012), these configurations were assessed using topic coherence, diversity, and qualitative interpretability.

Findings

Preprocessing Strategies

Analysis of the collected data showed distinct participation patterns: the U.S. dataset contained 46,107 unique users among 56,767 comments, whereas the South Korean dataset contained 22,557 among 41,279, indicating a higher concentration of heavy users in Korea. For length-based labeling, we used character counts rather than word counts because inconsistent spacing in Korean could distort token- or word-based estimates. Following information-theoretic studies showing that Korean syllables carry higher information density than English characters (Han et al. 1996; Pellegrino et al. 2011), percentile-based thresholds were adjusted by language.

Short comments were more common in the U.S. data and were excluded because they offered limited insight into meditation motivations. Extremely long comments were treated as emotional venting, more suitable for qualitative inquiry than automated modeling. The final corpus contained 40,905 U.S. documents and 24,169 Korean documents: 30,790 medium and 10,115 long comments in English, and 17,424 medium and 6,745 long comments in Korean.

Model Evaluation and Interpretability

As shown in Figure 1 and 2, the automated hierarchical aggregation in stage 3 achieved the highest coherence score, but compressed 19 topics into 4 groups, reducing thematic detail. This shows a key limitation of automated aggregation: although it efficiently merges topics close in lexical or embedding space, it may miss interpretive distinctions. By contrast, human-led classifications such as the Aarne–Thompson–Uther Index distinguish meaningful units such as tale types, motifs, actions, characters, and settings. Likewise, our human aggregation examined whether machine-generated topics should remain separate or be regrouped by contextual meaning.

For comparison, we evaluated automated and human-led aggregations at the same number of grouped topics. The human grouping was built through qualitative coding: experts identified eight themes, two independent coders assigned each topic to them, and intercoder reliability was measured. Under this condition, human aggregation showed higher coherence and lower diversity. This suggests that it did not arbitrarily reduce topics into themes, but systematically regrouped them by semantic distinctions that automated clustering could miss.

Figure Hierarchical clustering of the 19 topics used for automated topic aggregation.
Figure Topic model performance across data refinement and merging stages.

As supplementary validation, we also modeled each language corpus with language-specific embeddings: all-mpnet-base-v2 for English and snunlp/KR-SBERT for Korean. This allowed us to check whether comparable themes appeared across embedding strategies and whether the topics reflected culturally meaningful patterns rather than artifacts of multilingual embeddings.

Handling Unclassified Documents

Due to the informal and fragmented nature of YouTube comments (Thelwall 2018), many documents were assigned to the -1 category. We therefore re-examined them to separate noise from ambiguous or overlapping themes, modeled the outliers to detect residual thematic structure, and tested a classifier trained on examples from the eight themes to identify comments spanning multiple themes.

Contribution

This study contributes to Digital Humanities by examining how multilingual and multicultural corpora can be modeled, compared, and interpreted through BERTopic. It treats topic modeling not as a fully automated pipeline, but as an interpretive aid that requires continuous researcher intervention. To enable cross-linguistic comparison, the main analysis uses multilingual embeddings; to check possible English-centered bias, language-specific embeddings are used in sub-corpus analyses. The results show that linguistic structure and information density must be considered from preprocessing to evaluation. Most importantly, the study demonstrates that human-guided thematic aggregation can add interpretive value beyond automated clustering by distinguishing semantic differences among seemingly similar expressions and reorganizing topics according to contextual and cultural meaning. Finally, unclassified documents should not be dismissed as noise, since they may contain ambiguous, overlapping, or culturally nuanced discourse.

References
  1. Association of Internet Researchers [AoIR] (2020): “Ethical Guidelines for Internet Research”, Chicago: Association of Internet Researchers.
  2. Azher, I. A. / Seethi, V. D. R. / Akella, A. P. / Alhoori, H. (2024): “LimTopic: LLM-based Topic Modeling and Text Summarization for Analyzing Scientific Articles Limitations”, in: Proceedings of the 24th ACM/IEEE Joint Conference on Digital Libraries: 1–12.
  3. Berelson, B. (1952): Content Analysis in Communication Research. Glencoe, IL: Free Press.
  4. Bianchi, F. / Terragni, S. / Hovy, D. (2021): “Pre-training Is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence”, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Volume 2: Short Papers: 759–766.
  5. Blei, D. M. / Ng, A. Y. / Jordan, M. I. (2003): “Latent Dirichlet Allocation”, in: Journal of Machine Learning Research 3: 993–1022.
  6. Borčin, M. / Jose, J. M. (2024): “Optimizing BERTopic: Analysis and Reproducibility Study of Parameter Influences on Topic Modeling”, in: European Conference on Information Retrieval. Cham: Springer Nature Switzerland: 147–160.
  7. Braun, V. / Clarke, V. (2006): “Using thematic analysis in psychology”, in: Qualitative Research in Psychology 3, 2: 77–101.
  8. Chang, J. / Gerrish, S. / Wang, C. / Boyd-Graber, J. / Blei, D. (2009): “Reading Tea Leaves: How Humans Interpret Topic Models”, in: Advances in Neural Information Processing Systems 22.
  9. Cheng, K. / Inzer, S. / Leung, A. / Shen, X. / Perlmutter, M. / Lindstrom, M. / Needell, D. (2022): “Multi-scale Hybridized Topic Modeling: A Pipeline for Analyzing Unstructured Text Datasets via Topic Modeling”, in: arXiv preprint arXiv:2211.13496.
  10. Churchill, R. / Singh, L. (2022): “The Evolution of Topic Modeling”, in: ACM Computing Surveys 54, 10s: 1–35.
  11. Denny, M. J. / Spirling, A. (2018): “Text Preprocessing for Unsupervised Learning: Why It Matters, When It Misleads, and What to Do about It”, in: Political Analysis 26, 2: 168–189.
  12. Egger, R. / Yu, J. (2022): “A Topic Modeling Comparison Between LDA, NMF, Top2Vec, and BERTopic to Demystify Twitter Posts”, in: Frontiers in Sociology 7: 886498.
  13. Griffiths, T. L. / Steyvers, M. (2004): “Finding Scientific Topics”, in: Proceedings of the National Academy of Sciences 101, suppl. 1: 5228–5235.
  14. Grimmer, J. / Stewart, B. M. (2013): “Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts”, in: Political Analysis 21, 3: 267–297.
  15. Grootendorst, M. (2022): “BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure”, arXiv preprint arXiv:2203.05794 https://arxiv.org/abs/2203.05794 [05.05.2026].
  16. Han, Y. S. / Park, H. R. / Shin, J. H. / Choi, K. S. (1996): “An Upper Bound Estimate for the Entropy of Korean Texts”, in: Literary and Linguistic Computing 11, 3: 141–146.
  17. Kaur, A. / Wallace, J. R. (2024): “Moving beyond LDA: A Comparison of Unsupervised Topic Modelling Techniques for Qualitative Data Analysis of Online Communities”, arXiv preprint arXiv:2412.14486 https://arxiv.org/abs/2412.14486 [05.05.2026].
  18. Kozinets, R. V. (2002): “The Field behind the Screen: Using Netnography for Marketing Research in Online Communities”, in: Journal of Marketing Research 39, 1: 61–72.
  19. Marwick, A. E. / Boyd, D. (2011): “I Tweet Honestly, I Tweet Passionately: Twitter Users, Context Collapse, and the Imagined Audience”, in: New Media & Society 13, 1: 114–133.
  20. McInnes, L. / Healy, J. / & Astels, S. (2017): “hdbscan: Hierarchical Density Based Clustering”, in: Journal of Open Source Software 2, 11: 205.
  21. Meeks, E. / Weingart, S. B. (2012): “The Digital Humanities Contribution to Topic Modeling”, in: Journal of Digital Humanities 2, 1: 1–6.
  22. Mehrotra, R. / Sanner, S. / Buntine, W. / Xie, L. (2013): “Improving LDA Topic Models for Microblogs via Tweet Pooling and Automatic Labeling”, in: Proceedings of the 36th International ACM SIGIR Conference on Research and Development in Information Retrieval: 889–892.
  23. Papadimitriou, I. / Lopez, K. / Jurafsky, D. (2023): “Multilingual BERT has an Accent: Evaluating English Influences on Fluency in Multilingual Models”, in: Findings of the Association for Computational Linguistics: EACL 2023: 1194–1200.
  24. Pellegrino, F. / Coupé, C. / Marsico, E. (2011): “A Cross-Language Perspective on Speech Information Rate”, in: Language 87, 3: 539–558.
  25. Roberts, M. E. / Stewart, B. M. / Tingley, D. / Lucas, C. / Leder-Luis, J. / Gadarian, S. K. / Albertson, B. / Rand, D. G. (2014): “Structural Topic Models for Open-Ended Survey Responses”, in: American Journal of Political Science 58, 4: 1064–1082.
  26. Röder, M. / Both, A. / Hinneburg, A. (2015): “Exploring the Space of Topic Coherence Measures”, in: Proceedings of the Eighth ACM International Conference on Web Search and Data Mining: 399–408.
  27. Schofield, A. / Magnusson, M. / Mimno, D. (2017): “Pulling out the Stops: Rethinking Stopword Removal for Topic Models”, in: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, Volume 2: Short Papers: 432–436.
  28. Taddy, M. (2012): “On Estimation and Selection for Topic Models”, in: Artificial Intelligence and Statistics. PMLR: 1184–1193.
  29. Thelwall, M. (2018): “Social Media Analytics for YouTube Comments: Potential and Limitations”, in: International Journal of Social Research Methodology 21, 3: 303–316.
  30. Uther, H. (2004): The Types of International Folktales: A Classification and Bibliography. Based on the System of Antti Aarne and Stith Thompson. 3 vols. Helsinki: Suomalainen Tiedeakatemia.
  31. Wang, Q. / Ma, B. (2024): “Enhancing BERTopic with Pre-Clustered Knowledge: Reducing Feature Sparsity in Short Text Topic Modeling”, in: Journal of Data Analysis and Information Processing 12, 4: 597–611.