Daejeon, July 27–31
There is a long history of using data-driven methods to study and address lynchings in the United States. In the 1880s, the Chicago Tribune began publishing lynching records from across the nation. By the 1890s, Ida B. Wells was analyzing their data to dispel racist myths about the causes of lynchings (Wells, 1895). By the 1910s, the NAACP was also tabulating its own data and publishing sweeping reports (NAACP, 1919). So was Monroe Work at the Tuskegee Institute. All of these data-driven initiatives supported legal efforts and public campaigns, arming the anti-lynching movement with clear evidence about the scale and pervasiveness of the issue. These data remain important tools for addressing the painful legacy of lynchings in the United States. They record the places, victims, and details of thousands of cases over the course of decades.
This poster revisits the data first compiled by the anti-lynching movement. It does so by describing a newly constructed dataset: the Dataset of U.S. Lynching Reports (DUSLR). This dataset contains over 50,000 reports of lynchings identified in digitized newspapers from 1865 to 1921. These newspaper reports are relevant to the anti-lynching movement because newspapers were the means through which the movement compiled the vast majority of its data. These digitized newspaper reports–contemporaneous with the anti-lynching movement–subsequently reveal the information ecosystem that allowed the movement to enact its data-driven activism. In this way, DUSLR provides a new kind of resource for addressing lynchings. It is built upon the work of contemporaneous projects–the Chicago Tribune, Monroe Work, and the NAACP–as well as modern projects–specifically, the Tolnay-Beck-Bailey Inventory and the Seguin-Rigby Dataset–that document lynchings. But DUSLR differs from these records in that it aggregates more information about individual cases by identifying and compiling their contemporaneous newspaper reports. This dataset therefore offers new ways to address lynchings: it can be used to address individual cases or to study the larger discourse on lynchings in America as it appeared in newspapers.
This poster shows how DUSLR was constructed through the following computational process: 1) scraping the Chronicling America archive of digitized newspapers for the names of lynching victims recorded in lynching datasets, 2) creating loose clippings of the text surrounding victim names, 3) hand-keying a training and test set to be used to fine-tune BERT for binary classification (labelling clippings as either “references to a lynching” or “not references to a lynching”, 4) deploying our fine-tuned BERT-base model to filter and tag our scraped candidates, and 5) enriching the data with newspaper and case locations. The poster will also showcase DUSLR’s digital exhibit, allowing viewers to interact with the data. This digital exhibit includes interactive maps of newspaper and lynching locations as well as clipping images. It is designed specifically to highlight the close relationship between newspaper reports of lynchings and the data on lynchings compiled through the anti-lynching movement. Finally, because the data contains discourse on lynchings from a wide range of contemporaneous perspectives, many of which were racist and unapologetic about lynchings, the exhibit provides ethical frameworks for the data. Deploying what Milosev (2026) calls the Antidote Model for extremist corpora, the exhibit emphasizes that any knowledge claims about individual cases derived from DUSLR need to be critically analyzed and corroborated with other sources.
DUSLR and its digital exhibit offer important contributions to the digital humanities and the history of activism in the United States. This project provides new insight into the anti-lynching movement. It gives much due credit to the movement’s leaders, including Ida B. Wells, Monroe Work, and the NAACP–all of whom effectively applied data to their activist agenda. Through its use of text-mining and machine learning methods, the project also serves as an example of speculative bibliography: “an experimental approach to the digitized archive in which textual associations are constituted propositionally, iteratively, and (sometimes) temporarily, as the result of probabilistic computational models” (Cordell, 2022). The project is an innovative example because it uncovers and reconstitutes speculative versions of the newspaper networks that served as the foundations for nearly all lynching data. Finally, the project also provides a new digital exhibit that can be used by other researchers and scholars who wish to address the painful legacy of lynchings in the United States–a history which both deserves and requires more conscientious engagement.