DH 2026

Daejeon, July 27–31

Short Paper

The ‘K’ Factor: Comparing human and LLM-GeneratedEnglish Lyrics in K-pop

Andra-Maria Florescu
University of Bucharest, Romania · andra-maria.florescu@s.unibuc.ro
Anca Daniela Dinu
University of Bucharest, Romania · ancaddinu@gmail.com

Introduction

This paper is a continuation of the pilot study presented in Florescu (2025), comparing human written and AI-generated Kpop lyrics on English versions of Korean tracks and English original tracks.

Kpop is known for hybridization (Jin and Ryoo 2014) and for bending grammatical norms, especially of English (Willis 2014), as Kpop can sometimes be inspired by Western pop music (Li 2022; Barnes-Sadler et al. 2025).

Our contribution consists in introducing a manually curated dataset of human and AI-generated English Kpop song lyrics, and in comparing them by using both computational and manual close reading methods.

Related work

(Barnes-Sadler et al. 2025) focused on English usage within Korean language in Kpop tracks, rather than English-exclusive songs, treating mostly code-mixing or code-switching. (Frohmann et al. 2025) explored automated lyric generation and artistic assistance, without evaluating genre-specific stylistic fidelity or cultural nuance.

(Roe 2025; Klein et al. 2026) argued that generative AI outputs are cultural artifacts shaped by societal norms and biases, not neutral text, so they should be studied beyond their technical results. This directly applies to English lyric generation in Asian music contexts.

Although generative AI is heavily used in the music industry, few studies focus on their potential for song lyric production (Ding et al. 2025) and no research deals with English-exclusive Kpop songs, nor with LLMs’ ability to generate lyrics in such a specific context.

Data

The corpus comprises 230 songs, 200 human written and 30 AI-generated, totalling over 100,000 words. The Kpop lyrics authored by popular Korean artists were extracted from a web platform database

https://genius.com/

. To see how LLMs behave without additional instructions, we used a zero-shot prompt approach, with the prompt: "You are a successful Kpop idol and now want to produce a new song in English to be catchy and break records. Show the full song lyrics".

We also experimented with (creativity) parameters, (maximum temperature and top p) for models that had this option available. We included publicly available LLMs on their websites or from Huggingface

https://huggingface.co/spaces/

and DeepInfra

https://deepinfra.com/chat/

such as: ChatGPT, Claude, Copilot, Deepseek, Gemini, etc.

Results

Quantitative Analysis

We computed three basic metrics for both humans and machines: the average song length in words, the lexical diversity (unique words/total words), and the line repetition ratio (repeated lines/total lines), listed in Table 1. They reveal that human lyrics are longer, lexically richer, and contain more line repetitions (a typical K factor).

Language Style Matching score, which measures the use of function words, obtained with LIWC (Boyd et al. 2022), a tool designed to extract psycho-socio-linguistic categories, is rather high (0.82), confirming that AI matches the basic grammatical footprint of the human lyrics. Nevertheless, LLM-generated lyrics lack the structural complexity of Kpop, as shown by LIWC features given in Table 2. All the human scores are higher than the LLMs’ ones. Nonflu score, which represents the use of non-fluent grammar and wording, shows a spontaneity gap between Humans and LLMs, since human songs focus more on musical flow than on grammatical accuracy. The Filler scores (hooks like oh or uh) suggest that humans concentrate more on rhythmic padding to fit English words into fast-paced Korean melodies.

CategoryAvg words/songLexical diversityAvg line repetition ratio
LLMs3640.012830.2364
Humans4300.04370.3559

Table 1. Basic metrics

Humans also use Conversation (informal words like well, anyway), Netspeak (slang) and Assent (approval words like yeah, ok) more, denoting lyrics written for an audience, to create the typical Kpop "idol-style energy". The AI songs are more clean, even sterile and seem over-done. By stripping away the noise or energy, the LLMs fail to replicate the adlibs and rhythmic style that define human Kpop performance.

CategoryConversationNetspeakAssentNonfluFiller
LLMs1.960.90.890.120.05
Humans5.612.061.541.350.76

Table 2. LIWC scores

Sentiment Analysis

To extract the sentiments from the lyrics we used Hartmann’s seven class emotion model DistilRoberta

https://huggingface.co/j-hartmann/emotion-english-distilroberta-base

. The results show that LLMs tend to be more positive than the humans (86.7% vs. 57%, p < 0.002, d = 0.76), who cover a wider range of emotions, darker and more ambiguous content, more disgust and sadness-coded language.

Readability and Perplexity

Readability metrics indicated that the LLMs produce higher grade level lyrics, use longer words, more elevated vocabulary, and adhere to grammar rules, whereas humans employ wordplay and genre-specific deviations from English conventions, as noticed by (Schneider 2024), probably since human lyrics are optimized for auditory processing, while the LLMs tend to mimic written poetry.

Perplexity scores computed with GPT-Neo

https://huggingface.co/EleutherAI/gpt-neo-1.3B

revealed that the human lyrics are more in line with the genre and vary less (mean ≈ 6.8, SD ≈ 2.6) than LLMs’ lyrics which are more unnatural (mean ≈ 9.5, SD ≈ 8.3).

Automatic Classification

Since the dataset is imbalanced, we built a frozen Distilbert

https://huggingface.co/docs/transformers/en/model doc/distilbert

classifier with in-fold augmentation and SMOTE

https://imbalanced-learn.org/stable/references/generated/imblearn.over sampling.SMOTE.html

, achieving an F1 score of 0.955 (precision=1.00, recall=0.917) across 10 repeated 60/40 splits with zero false positives, and three false negatives.

Empirical Observations

We noticed that some models generated lyrics resembling more with generic pop instead of having "K" elements. LLMs used figurative language and elevated vocabulary but lacked the lyricism patterns characteristic of successful English Kpop songs.

Conclusion

LLMs demonstrated a moderate ability to generate English Kpop lyrics and struggled to authentically replicate the distinctive stylistic and cultural nuances of the genre. Their outputs either exaggerated certain lyrical tropes or underperformed in conveying the typical Kpop emotional depth, thematic consistency, rhythm, and playful wordplay.

Not only LLM-generated lyrics are of lower quality, but the overuse of generative AI in the creative industries poses major environmental sustainability risks. This technology requires massive amounts of resources, high carbon emissions which can negatively impact the climate, as well as society (Esho et al. 2026).

Acknowledgments

This work was supported by a grant of the Ministry of Education and Research, CCCDI - UEFISCDI, project number PN-IV-P6-6.1-CoEx-2024-0042, within PNCDI IV, and by a grant of the Ministry of Research, Innovation and Digitization, CNCS - UEFISCDI, project SIROLA, number PN-IV-P1-PCE-2023-1701, within PNCDI IV.

References
  1. Barnes-Sadler, Simon / Ahn, Hyejeong / Kiaer, Jieun (2025). A quantitative characterisation of english in k-pop through the korean wave 2000–2020. Asian Englishes, 27(3):700–715, 2025. doi: 10.1080/13488678.2025.2533537.
  2. Boyd, Ryan L. / Ashokkumar, Ashwini / Seraj, Sarah / Pennebaker, James W (2022): The development and psychometric properties of liwc-22. Austin, TX: University of Texas at Austin.
  3. Ding, Shuangrui / Liu, Zihan / Dong, Xiaoyi / Zhang, Pan / Qian, Rui / Huang, Junhao / He, Conghui / Lin, Dahua / Wang, Jiaqi (2025): “SongComposer: A large language model for lyric and melody generation in song composition”, in: Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.): Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7108–7127, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long.352.
  4. Esho, Esther O. / Akinyelu, Andronicus A. / Dinis, Maria Alzira Pimenta (2026): “Sustainable generative ai and quantum computing: review assessment on the environmental impact of generative ai and quantum technologies”, in: Frontiers in Sustainability, Volume 7 - 2026, 2026. ISSN 2673-4524. doi: 10.3389/frsus.2026.1726832.
  5. Florescu, Andra-Maria (2025) “Llms and kpop: Can llms be good lyricists?” in: Book of Abstracts: 3rd International Conference on Recent Advances in Digital Humanities (RADH 2025), page 17, Craiova, Romania.
  6. Frohmann, Markus / Epure, Elena V. / Meseguer-Brocal, Gabriel / Schedl, Markus / Hennequin, Romain (2025): “Ai-generated song detection via lyrics transcripts” <https://arxiv.org/abs/2506.18488>.
  7. Jin, Dal Yong / Ryoo, Woongjae (2014): “Critical interpretation of hybrid k-pop: The global-local paradigm of english mixing in lyrics” in: Popular Music and Society, 37(2):113–131.
  8. Klein, Lauren / Martin, Meredith / Brock, André / Antoniak, Maria / Walsh, Melanie / Johnson, Jessica Marie / Tilton, Lauren / Mimno, David (2026): “Provocations from the humanities for generative ai research” <https://arxiv.org/abs/2502.19190>.
  9. Li, Xingnuo (2022): “Reasons for the success of kpop (korean popular music) culture in the international spread” in: Proceedings of the 2022 8th International Conference on Humanities and Social Science Research (ICHSSR 2022), pages 2617–2621. Atlantis Press, ISBN 978-94-6239-580-0. doi: 10.2991/assehr.k.220504.475.
  10. Roe, Jasper (2025): “Generative ai as cultural artifact: Applying anthropological methods to ai literacy” in: Postdigital Science and Education, 7(4):1107–1124, 2025. ISSN 2524-4868. doi: 10.1007/s42438-025-00547-y.
  11. Schneider, Ian (2024): “English’s expanding linguistic foothold in k-pop lyrics: A mixed methods approach” in: English Today, 40(2):105–112, 2024. doi:10.1017/S0266078423000275.
  12. Willis, Jain (2014): “Putting the k in k-pop: Korean or konglish pop music?” in: Word of Mouth, BYU, Winter, 2014:1–12