DH 2026

Daejeon, July 27–31

Poster

Forgotten Stories about Invisible Empire:A Large-scale Text Analysis of the Klan Stories in US Historical Newspaper Archives

Jaihyun Park
Indiana University Indianapolis, United States of America · jaipark@iu.edu

<Warning: This article includes examples of words targeting marginalized populations.>

Archives as a memory institution are where power struggle happens (Yale, 2015), and the digitization of history is disguised as technology (objective and value-neutral), but in fact, curating what to digitize is a political process of what to include, which also can be translated to what to exclude (Zaagsma, 2023). While there have been many studies on the history of the Ku Klux Klan (KKK) (Cunningham, 2013; McVeigh, 2009), understanding the stories about the KKK remains largely unexplored through distant reading

While recognizing Franco Moretti’s significant influence on digital humanities with the concept of distant reading, author(s) also acknowledge the serious allegations of sexual misconduct against him and express sincere sympathy for those harmed. For critical reflections and community responses within digital humanities, see: https://lklein.com/digital-humanities/distant-reading-after-moretti/

approaches (Underwood, 2016). Studying how an organization known for extrajudicial violence is portrayed in digital archives gives us an opportunity to understand the ways in which digitization mediates historical memory – transforming acts of brutality into sanitized narratives that mystify their true nature. This study adds to the scholarship on empirical analysis of digital archives, especially the historical divide in the US (Park & Cordell, 2023; Park & Cordell, 2025).

By focusing on the second period of the KKK organization, from 1915 to 1944, this project collected 375,444 articles by using the keywords

Keywords include the organization itself: “klan”, “khan” (as a possible OCR error for “klan”), “klux”, “kloran”, “klavern”, “klankraft”, “klandom”, “konvention”, “klonverse”, “invisible empire”, “imperial wizard”, “grand dragon”, “knight rider”, “night rider”, “nightrider” and violent activities that might have involved KKK: “mask mob”, “cross burn”, “flog”, “lynch”, “vigilante”.

from the American Stories dataset (Dell et al., 2023). To consider possible false positive cases, which might have included the word “lynch” as the last name of the person, Named Entity Recognition (NER) (Qi et al., 2020) was applied and filtered out when the word “lynch” was used in the names. On the final 115,645 data, the embedding-based topic modeling (Angelov & Inkpen, 2024), the “distiluse-base-multilingual-cased” was applied to create an embedding for each article, and topic vectors are jointly embedded in the vector space. A total of 546 topics have been identified, and the article embeddings, colored by topic number, are visualized in a three-dimensional space using UMAP dimension reduction (See Figure 1).

<Figure 1: The vector representation of KKK articles with UMAP dimension reduction.
The data points represent articles, and the color represents topics.>The data points represent articles, and the color represents topics.>

034515
negroprosecutionarrestsnegrosentencing
negroesjudicialpolicianegroespenal
blackssentencingarrestblacksjudicial
arrestsprosecutionsarrestingracialjustice
blackwellacquittedpolicemanniggerlegislative
arrestprosecutorpolicemenniggerslegislature
arrestingjurypoliceblackimpunity
niggerconvictedsheriffblackwellprosecute
niggerslitigationcriminalsslaveryjail
policiatribunaljailedinterracialpenitentiary

<Table 1: Top 5 relevant topics except identified topics due to OCR error
or the false positive article segmentation>or the false positive article segmentation>

The most frequent topic identified to be relevant to the discourse around the KKK (See Table 1) is about African Americans, the population who had been the victims of extrajudicial violence (Topic 0: 486 articles), and Topic 5 (336 articles) also includes racial words. Articles include “Thomas R. Marshall a telegram de-manding the dispatch of federal troops to Mississippi TO protect citizens from anarchy and mob violence" following the lynching and burning of John Hart field. negro at Ellisville. Miss. yes terday.” (Topic 0, Topic Score: 0.8585) from The Chattanooga News, Tennessee, on June 27th, 1919 and “Butte, Mont, Aug. j2.-Citizens spent a restless night, owing to the rumors of holesale lynchings and threatened outbreaks by comrades OF Frank Little, the Industrial Worker of the World leader who was lynch ed yesterday” (Topic 5, Topic Score: 0.7935) from the Brunswick News on August 3rd, 1917.

Topic 3 (378 articles) and Topic 15 (243 articles) include topic words about the justice system. An article from The Ocala Evening Star, Florida on May 8th, 1918 has been clustered to Topic 3; “If the Alabama klu klux intend to stim- ulate into industry white as well as colored loafers, and if they have no gentlemen of leisure in their own ranks, the Star will approve of them” (Topic Score: 0.7864). Another article “The lynching of Robert P. Praeger, a native OF Germany, at Collinsville, Illinois, Thursday night is deplored by all right-thinking people, So far as can be learned, there was nothing in the mans words or acts to cause a gang of irresponsible men to take him from the regularly constituted law officers and murder him.” was published on April 6th, 1918 by Alexandria Gazette, Washington D.C. and classified as Topic 5 (Topic Score: 0.7640).

As seen from the example of Topic 3 and Topic 15, the result of topic modeling does not clearly differentiate the discourse of endorsing the KKK from the discourse of rejecting the extrajudicial violence. Both examples are clustered under a similar topic. While this study remains a preliminary study, the future direction will include developing a pipeline to further understand competing discourse around the KKK.

References
  1. Angelov, D., & Inkpen, D. (2024). Topic modeling: Contextual token embeddings are all you need. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 13528-13539).
  2. Cunningham, D. (2013). Klansville, USA: The rise and fall of the civil rights-era Ku Klux Klan. Oxford University Press.
  3. Dell, M., Carlson, J., Bryan, T., Silcock, E., Arora, A., Shen, Z., Amico-Wong, L. D., Quan, L., Querbin, P., & Heldring, L. (2023). American stories: A large-scale structured text dataset of historical us newspapers. Advances in Neural Information Processing Systems36, 80744-80772.
  4. McVeigh, R. (2009). The rise of the Ku Klux Klan: Right-wing movements and national politics (Vol. 32). U of Minnesota Press.
  5. Park, J., & Cordell, R. (2023). A Quantitative Discourse Analysis of Asian Workers in the US Historical Newspapers. In The Joint 3rd International Conference on Natural Language Processing for Digital Humanities and 8th International Workshop on Computational Linguistics for Uralic Languages (pp. 7-15).
  6. Park, J., & Cordell, R. (2025). A Data-driven Investigation of Euphemistic Language: Comparing the usage of “slave” and “servant” in 19th century US newspapers. In The Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities (NLP4DH), (pp. 350-364).
  7. Underwood, T. (2016). Distant reading and recent intellectual history. Debates in the Digital Humanities2016, 530-33.
  8. Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python natural language processing toolkit for many human languages. In The Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL): System Demonstrations. (pp.101-108).
  9. Yale, E. (2015). The history of archives: the state of the discipline. Book History18(1), 332-359.
  10. Zaagsma, G. (2023). Digital history and the politics of digitization. Digital Scholarship in the Humanities38(2), 830-851.