DH 2026

Daejeon, July 27–31

Fri, July 3114:00–15:30S073209-211
Short Paper

Engaging Historians in the Archived Web: A plug-in for SolrWayback

Victor Harbo Johnston
Aarhus University, Denmark · vijo@cas.au.dk
Helle Strandgaard Jensen
Aarhus University, Denmark · hs.jensen@cas.au.dk
Jørn Thøgersen
Aarhus University, Denmark · jorn@cas.au.dk

Introduction

Historians are traditionally trained to do so-called close reading analysis. They carefully collect and read a few documents at a time, a method that gives them extensive and detailed knowledge of all the sources they are working with. So, what do historians do when confronted with half a billion sources in multiple languages? That is a fundamental methodological question we face every day in the WEB CHILD project.

In WEB CHILD we investigate how the introduction of the WWW has changed childhood in the United States, Denmark and South Korea. We ask traditional questions about childhood history but use a ~50TB corpus of archived web pages to answer them.

One of our key solutions is to develop a plug-in for the open-source software SolrWayback (Egense et al., 2026). The plug-in will help bridge the gap between historians’ traditional ways of thinking and working and the vastness and complexity of the archived web. The solution is a two-pronged approach facilitating analysis on small and big scale. (Egense et al., 2026). The plug-in will help bridge the gap between historians’ traditional ways of thinking and working and the vastness and complexity of the archived web. The solution is a two-pronged approach facilitating analysis on small and big scale.

In this paper, we present our plug-in which supports historians and other scholars interested in using archived web material in their analyses.

Background

Contemporary historians shun using archived versions of the Web as a source in historical studies; a challenge yet to be addressed as has been pointed out repeatedly in the past years (Brügger, 2021; Brügger et al., 2017; Cole, 2011; Milligan, 2019; Schafer & Thierry, 2018; Webster, 2017; Winters, 2018).

There have been several attempts to make web archives more accessible to historians and other scholars within the humanities (Ruest et al., 2022; Sherratt et al., n.d.) However, judging from the citations on Google Scholar, one of the most successful projects, Archives Unleashed, it is not historians but librarians and digital humanities scholars who primarily use these tools. One of the reasons for the lack of engagement with these tools by traditional humanities scholars might be the digital complexity of the archival format, WARC (Ruest et al., 2022; Sherratt et al., n.d.) However, judging from the citations on Google Scholar, one of the most successful projects, Archives Unleashed, it is not historians but librarians and digital humanities scholars who primarily use these tools. One of the reasons for the lack of engagement with these tools by traditional humanities scholars might be the digital complexity of the archival format, WARC (Maemura, 2023). Another might be related to one of the overall ontological debates in digital humanities between these kinds of data-driven tools and more traditional problem-driven approaches (Maemura, 2023). Another might be related to one of the overall ontological debates in digital humanities between these kinds of data-driven tools and more traditional problem-driven approaches (Eunice Gonzalez & Vitti Rodrigues, 2022; Heinsen, 2023).(Eunice Gonzalez & Vitti Rodrigues, 2022; Heinsen, 2023).

Plug-in solution developed in the WEB CHILD project

In the WEB CHILD team, we are building a plug-in which allows users to work in problem-driven ways applying traditional methods from source criticism and combine these with distant reading of large collections. We do this in several ways, two examples are:

For the small-scale, close reading we provide an interface which solves the problem of hidden temporal differences in the browser rendering of the single sources from the archived webpages, the so-called playback view. One of the biggest problems for understanding the archived webpage as a source is the time difference in the rendering, where items (pictures, text, styling) are often from different dates, but are presented as a coherent entity from one specific date. Our solution makes the user aware of this time inconsistency when navigating the singular sources in the playback view; the modus operandi most likely applied when working in a problem-driven way.

For the big-scale distant reading we tackle the problem of distinguishing between archival date and production date of the sources. One of the biggest challenges in using web archives as a historian is the ways in which the material is collected. When institutions archive the web they harvest, they only provide the time of harvest and not the time of production. One way of fixing this is to create a signal fusing approach, that looks for signals of the production date in the archived material.

SolrWayback plug-in: an open-source solution

Tools for discovery and playback of the archived web do exist; the most well-known is probably the American Internet Archive’s Wayback Machine, which provides URL-based playback of archived webpages. Our plug-in is built around another piece of software for playback of material from the archived web: the open-source software SolrWayback. This software provides full-text search and multiple ways of interacting with the archived material (Kurzmeier, 2025).(Kurzmeier, 2025).

Our choice of SolrWayback aligns with the ideas of reusability, sustainability and enhanced engagement with the open-source community. We are already working with the Royal Danish Library on long term incorporation of the plug-in into their infrastructure. Our approach to data sustainability focuses on using international standards such as the WARC format.

Conclusion

Modern cultural heritage is increasingly digital in nature. Our plug-in helps humanities scholars overcome the problems they face when working the archived web. In terms of the wider digital humanities community, working with archived web presents a new sub-field which until now has mostly been confined to practitioners of digital archiving. In doing so we contribute to the diversity of the DH community.

References
  1. Brügger, N. (2021). Digital humanities and web archives: Possible new paths for combining datasets. In International Journal of Digital Humanities (Vol. 2, Issue 1, pp. 145–168). https://doi.org/10.1007/s42803-021-00038-z
  2. Brügger, N., Laursen, D., & Nielsen, J. (2017). Exploring the domain names of the Danish web. In N. Brügger & R. Schroeder (Eds.), The Web as History (pp. 62–80). UCL Press. http://www.jstor.org/stable/j.ctt1mtz55k.9
  3. Cole, J. (2011). Blogging Current Affairs History. Journal of Contemporary History, 46(3), 658–670. https://doi.org/10.1177/0022009411403341
  4. Egense, T., Eskildsen, T., Lauridsen, J., O’Brien, B., Jackson, A., Bellony, L., Thøgersen, J., & Johnston, V. H. (2026). SolrWayback (Version 5.4.1) [Computer software]. https://doi.org/10.5281/zenodo.18314399
  5. Eunice Gonzalez, M., & Vitti Rodrigues, M. (2022). Digital Humanities: Ethical Implications and Interdisciplinary Challenges. Humanities Bulletin, 5(1), 111–125.
  6. Heinsen, J. (2023). Kilde og data: Overvejelser om historiefaget og de digitale metoder. Temp, (26).
  7. Kurzmeier, M. (2025). Contextualizing and unlocking political web defacements for research. Journal of Digital History, (preprint).
  8. Maemura, E. (2023). All WARC and no playback: The materialities of data-centered web archives research. Big Data & Society, 10(1), 20539517231163172. https://doi.org/10.1177/20539517231163172
  9. Milligan, I. (2019). History in the Age of Abundance?: How the Web Is Transforming Historical Research. McGill-Queen’s University Press.
  10. Ruest, N., Fritz, S., & Milligan, I. (2022). Creating order from the mess: Web archive derivative datasets and notebooks. Archives and Records, 43(3), 316–331. https://doi.org/10.1080/23257962.2022.2100336
  11. Schafer, V., & Thierry, B. (2018). Web History in Context. In N. Brügger & I. Milligan (Eds.), The SAGE Handbook of Web History (pp. 59–72). SAGE.
  12. Sherratt, T., Jackson, A., & Bickford, J. (n.d.). GLAM-Workbench/web-archives (Version 1.2.0) [Computer software]. Retrieved https://doi.org/10.5281/zenodo.7898218
  13. Webster, P. (2017). DIGITAL CONTEMPORARY HISTORY SOURCES, TOOLS, METHODS, ISSUES. Temp - tidsskrift for historie, 7(14), Article 14.
  14. Winters, J. (2018). Web Archives and (Digital) History: A Troubled Past and a Promising Future? In N. Brügger & I. Milligan (Eds.), The SAGE Handbook of Web History (pp. 593–605). SAGE.