Daejeon, July 27–31
Introduction
Historians are traditionally trained to do so-called close reading analysis. They carefully collect and read a few documents at a time, a method that gives them extensive and detailed knowledge of all the sources they are working with. So, what do historians do when confronted with half a billion sources in multiple languages? That is a fundamental methodological question we face every day in the WEB CHILD project.
In WEB CHILD we investigate how the introduction of the WWW has changed childhood in the United States, Denmark and South Korea. We ask traditional questions about childhood history but use a ~50TB corpus of archived web pages to answer them.
One of our key solutions is to develop a plug-in for the open-source software SolrWayback (Egense et al., 2026). The plug-in will help bridge the gap between historians’ traditional ways of thinking and working and the vastness and complexity of the archived web. The solution is a two-pronged approach facilitating analysis on small and big scale. (Egense et al., 2026). The plug-in will help bridge the gap between historians’ traditional ways of thinking and working and the vastness and complexity of the archived web. The solution is a two-pronged approach facilitating analysis on small and big scale.
In this paper, we present our plug-in which supports historians and other scholars interested in using archived web material in their analyses.
Background
Contemporary historians shun using archived versions of the Web as a source in historical studies; a challenge yet to be addressed as has been pointed out repeatedly in the past years (Brügger, 2021; Brügger et al., 2017; Cole, 2011; Milligan, 2019; Schafer & Thierry, 2018; Webster, 2017; Winters, 2018).
There have been several attempts to make web archives more accessible to historians and other scholars within the humanities (Ruest et al., 2022; Sherratt et al., n.d.) However, judging from the citations on Google Scholar, one of the most successful projects, Archives Unleashed, it is not historians but librarians and digital humanities scholars who primarily use these tools. One of the reasons for the lack of engagement with these tools by traditional humanities scholars might be the digital complexity of the archival format, WARC (Ruest et al., 2022; Sherratt et al., n.d.) However, judging from the citations on Google Scholar, one of the most successful projects, Archives Unleashed, it is not historians but librarians and digital humanities scholars who primarily use these tools. One of the reasons for the lack of engagement with these tools by traditional humanities scholars might be the digital complexity of the archival format, WARC (Maemura, 2023). Another might be related to one of the overall ontological debates in digital humanities between these kinds of data-driven tools and more traditional problem-driven approaches (Maemura, 2023). Another might be related to one of the overall ontological debates in digital humanities between these kinds of data-driven tools and more traditional problem-driven approaches (Eunice Gonzalez & Vitti Rodrigues, 2022; Heinsen, 2023).(Eunice Gonzalez & Vitti Rodrigues, 2022; Heinsen, 2023).
Plug-in solution developed in the WEB CHILD project
In the WEB CHILD team, we are building a plug-in which allows users to work in problem-driven ways applying traditional methods from source criticism and combine these with distant reading of large collections. We do this in several ways, two examples are:
For the small-scale, close reading we provide an interface which solves the problem of hidden temporal differences in the browser rendering of the single sources from the archived webpages, the so-called playback view. One of the biggest problems for understanding the archived webpage as a source is the time difference in the rendering, where items (pictures, text, styling) are often from different dates, but are presented as a coherent entity from one specific date. Our solution makes the user aware of this time inconsistency when navigating the singular sources in the playback view; the modus operandi most likely applied when working in a problem-driven way.
For the big-scale distant reading we tackle the problem of distinguishing between archival date and production date of the sources. One of the biggest challenges in using web archives as a historian is the ways in which the material is collected. When institutions archive the web they harvest, they only provide the time of harvest and not the time of production. One way of fixing this is to create a signal fusing approach, that looks for signals of the production date in the archived material.
SolrWayback plug-in: an open-source solution
Tools for discovery and playback of the archived web do exist; the most well-known is probably the American Internet Archive’s Wayback Machine, which provides URL-based playback of archived webpages. Our plug-in is built around another piece of software for playback of material from the archived web: the open-source software SolrWayback. This software provides full-text search and multiple ways of interacting with the archived material (Kurzmeier, 2025).(Kurzmeier, 2025).
Our choice of SolrWayback aligns with the ideas of reusability, sustainability and enhanced engagement with the open-source community. We are already working with the Royal Danish Library on long term incorporation of the plug-in into their infrastructure. Our approach to data sustainability focuses on using international standards such as the WARC format.
Conclusion
Modern cultural heritage is increasingly digital in nature. Our plug-in helps humanities scholars overcome the problems they face when working the archived web. In terms of the wider digital humanities community, working with archived web presents a new sub-field which until now has mostly been confined to practitioners of digital archiving. In doing so we contribute to the diversity of the DH community.