DH 2026

Daejeon, July 27–31

Thu, July 3013:40–15:10S057209-211
Short Paper

A never-ending story: supporting access to museum and library collections

Mia Ridge
British Library, United Kingdom; Museum Data Service, United Kingdom · mia.ridge@bl.uk
Arran Rees
Museum Data Service, United Kingdom · arran@collectionstrust.org.uk
Ross Parry
University of Leicester, United Kingdom · ross.parry@leicester.ac.uk

In 2016, the Culture White Paper set out the UK government's 'ambition and strategy for the cultural sectors'. It included a vision of users enjoying 'a seamless experience online', with the ability 'to access particular collections in depth as well as search across all collections (Department for Digital, Culture, Media & Sport 2016)'.

In 2026, users have many options for accessing UK cultural heritage collections online, but the experience is far from seamless. Infrastructure for digital cultural heritage is a patchwork of open access images on non-profit platforms, paywalls and commercial databases, research datasets and various forms of 'free to view' collections - some collections can be accessed in depth, others only as metadata or catalogue records, while others have no online presence (Gosling et al. 2022). Searching across all collections is impossible.

Focusing on access to UK museum and library collections, this paper discusses two significant components of the UK's digital cultural heritage access for researchers.

The first is an aggregator of museum collections data, while the second is a brief description of the many options for accessing digitised items from the British Library. It closes with a call to support the long-term preservation and sustainability of cultural heritage collections online.

Aggregating museum collection data: introducing the Museum Data Service

The museum sector in the UK, as in many regions, is highly diverse. Large, national institutions can publish collection records with images, rich metadata and more via specialist websites and APIs, while tiny local or specialist museums might not be able to easily access their own collections records. While some of the estimated 80 million catalogue records held in c. 1700 accredited museums are available online, they were difficult to find and lacked persistent identifiers (Gosling et al. 2022). Finding museum records required using hundreds of websites with dozens of different interfaces, and collating museum collections for research at scale was almost impossible.

To address this challenge, the Museum Data Service (MDS) was launched in September 2024 to connect and share all the object records across all UK museums, large and small.

The MDS museumdata.uk was officially launched by Art UK, Collections Trust and the University of Leicester with start-up funding from Bloomberg Philanthropies. It is now part of the UK Arts and Humanities Research Council's portfolio of digital infrastructures.

The MDS offers UK museums, and institutions with museum-like collections, the ability to share their object records with other museums, researchers and the public, on their own terms.

The MDS platform responds to the sector's difficulties using earlier aggregation platforms (such as Culture Grid and Europeana), particularly the difficulties museums had in exporting collections data and formatting it for aggregation. Sharing records with Europeana requires organisations to prepare their data by mapping their records to the Europeana Data Model (Europeana PRO, n.d.). In contrast, the MDS works with each contributing museum's own levels of technical capacity - it acknowledges the reality of the inconsistent approaches to data standards across the sector and accepts data in any format. The MDS team maps incoming fields to align them with the ‘Spectrum Units of Information’ names, without reformatting or standardising the data.

Spectrum is a 'museum collections management standard' that originated in the UK. (‘Spectrum’, n.d.)

Museums can choose an appropriate Creative Commons licence for their dataset, reflecting their own access policies and attitude to risk management; this is also designed to encourage wider uptake by museums.

The real benefit of the MDS for Digital Humanities researchers is the meaningful connections it enables between museum collections and researchers (Parry et al. 2025). The MDS lets anyone search for types of collections, or specific types of objects within hundreds of museum collections. A user-friendly search interface enables detailed searches and sets of search results can be exported to support ongoing research by individuals. A specialist API supports the creation of bespoke datasets to support data analysis and modeling in heritage studies, for example to study questions of bias in cataloguing.

The fact that MDS does not ask collections to use a standard schema reduces the barriers to participation for museums. This also means that researchers can have access to the unique raw records that make up museum data, rather than a 'cleaned' version (Rawson and Muñoz 2016). The MDS does not make assumptions about the types of data that researchers might want, and it does not attempt to meet every need in the heritage sector. Rather, it aims to be an intrinsic part of a wider ecosystem for museum and heritage information, including participation in academic research projects (Bailey-Ross et al. 2024; Bailey et al. 2024). However, the MDS acknowledges that some researchers would prefer a slightly more standardised version of the data and is exploring options for this.

We hope to inspire future research using the museum data available on the platform, and to understand how the MDS might be developed to support emerging computational methods and Digital Humanities research questions.

The paradox of access enabled by commercial digitisation

While many nations have funded the digitisation of their national collections, organisations in the UK have been required by successive governments to seek public-private partnerships to digitise their collections. Under these deals, commercial companies who pay for collections digitisation can charge for access to those collections for a certain number of years. While research or direct government funding has paid for a certain amount of digitisation in the past, it is dwarfed by commercial funding. As a blog post by a British Library curator puts it:

'It has long been the goal of the British Library to make some of its digitised newspapers freely available online, but we also want to see the BNA [British Newspaper Archive] succeed as it has been doing, without which we could not have reached such a huge collection overall of digitised newspapers, nor the rate at which they are being produced (currently around half a million pages are being added to the BNA every month).' (McKernan 2021)

While most material digitised by commercial companies is eventually made more freely available, paywalls are one important factor in the fragmentation of access to collections. This was extensively documented in a Digital Collections Audit commissioned for the AHRC-funded research programme Towards a National Collection (TaNC) (Gosling et al. 2022).

The variety of sources of access to digitised collections from the British Library as in September 2023 helps illustrate this point.

In October 2023, the British Library suffered an extensive ransomware attack. (British Library 2024)

The British Library's websites and catalogues hosted hundreds of thousands digitised books, manuscripts, 3D objects and sound files. Open access images and metadata were also available on Flickr Commons, Europeana and Wikimedia Commons. Newspapers were available on the British Newspaper Archive. Other commercial databases held selected other items, while other items were available via union catalogues. The Library's Research Repository had more open access items, from digitised images to datasets created via research projects like Living with Machines.

https://bl.iro.bl.uk/

A large backlog of items was being processed for publication online. Altogether, this made answering an apparently simple enquiry about the availability of a particular item or collection online surprisingly complex.

However, this patchwork of platforms had one positive effect. While library users lost access to online catalogues, digitised and born-digital collections and the many services hosted by the British Library in a ransomware attack in October 2023, collections hosted on commercial and non-profit third-party platforms (including the Research Repository) continued to be available. The saying 'Lots of Copies Keep Stuff Safe' has never been more pertinent.

At the time of writing, the UK government is again investing in digital infrastructure, including the TaNC project and its successor, N-RICH. N-RICH has commissioned a range of activities to 'examine the scope, costs, risks, impacts and benefits of a future digital research infrastructure for cultural heritage in the UK' (Towards a National Collection 2025).

Digital records are fragile: the need to fund long-term access

Enabling access to collections online is one thing. Sustaining that access in the longer-term is another. The ransomware attack that took out most of the British Library's services was just the most dramatic incident that reduced online access to collections. Websites and platforms gradually disappearing when funding ends has eroded earlier digitisation efforts (Dunning 2009). For example, the results of the first big digitisation effort in the UK, the New Opportunities Fund c. 2000 - 2004, have long since been lost. Culture Grid, the UK's aggregator for Europeana, itself built on an earlier no-longer-supported project (the Peoples Network Discover Service), failed to find a sustainable financial model in the 2010s (Gosling 2019; Poole 2015).

More recently, collections sites have struggled to stay online through surges of bots scraping their sites to feed AI models. For example, the British Library's Research Repository was being scraped so heavily that it effectively experienced a Denial of Service attack until the team were able to put preventative measures in place.

Resourcing digital preservation and sustainability is not easy. In order to make the MDS a relatively lean operation, an early decision was made to exclude images and other media from the data aggregated. In addition to reducing financial costs, this reduces the environmental overhead of the service. However, not only does this add additional retrieval steps for researchers wanting to access images, it also means that the MDS cannot act as a backup of last resort for museums who lose access to their image assets.

Looking back over the long history of digitisation in museums, libraries and archives in the UK, it is clear that new economic models are needed to ensure the long-term availability of digital cultural heritage collections. Coming up with these models will require creativity and persuasive powers worthy of the collections they would protect.

References
  1. Bailey, Rebecca, Javier Pereda, Chris Michaels, and Tom Callahan. 2024. Unlocking the Potential of Digital Collections. A Call to Action. Arts and Humanities Research Council. https://doi.org/10.5281/zenodo.13838916.
  2. Bailey-Ross, Claire, Emily Burgess, and Panagiotis Papageorgiou. 2024. User Research: UK Gallery, Library, Archive and Museum  (GLAM) Digital Collections Infrastructure. Towards a National Collection. https://doi.org/10.5281/zenodo.12751226.
  3. British Library. 2024. Learning Lessons from the Cyber-Attack: British Library Cyber Incident Review. British Library. https://www.bl.uk/home/british-library-cyber-incident-review-8-march-2024.pdf/.
  4. Department for Digital, Culture, Media & Sport. 2016. The Culture White Paper. Department for Digital, Culture, Media & Sport. https://www.gov.uk/government/publications/culture-white-paper.
  5. Dunning, Alastair. 2009. ‘Digitising the Past: Next Steps for Public-Sector Digitisation’. Digital Information-Order Or Anarchy? http://eprints.rclis.org/handle/10760/18048.
  6. Europeana PRO. n.d. ‘Europeana Data Model’. Accessed 7 May 2026. https://pro.europeana.eu/page/edm-documentation.
  7. Gosling, Kevin. 2019. ‘Nothing New except What Has Been Forgotten’. Collections Trust, November 27. https://collectionstrust.org.uk/blog/nothing-new-except-what-has-been-forgotten/.
  8. Gosling, Kevin, Gordon McKenna, and Adrian Cooper. 2022. Digital Collections Audit. Towards a National Collection. Collections Trust. https://doi.org/10.5281/zenodo.6379581.
  9. McKernan, Luke. 2021. ‘Free to View Online Newspapers’. The Newsroom Blog, August 9. https://blogs-archive.bl.uk/thenewsroom/2021/08/free-to-view-online-newspapers.html.
  10. Parry, Ross, Stef De Sabbata, Andrew Ellis, Helen Hardy, and Mia Ridge. 2025. ‘A Framework for Sustainable and Scalable Cultural Data Integration and Analysis: Using the Large Dataset of the Museum Data Service’. 2025 11th International Symposium on System Security, Safety, and Reliability (ISSSR), April, 417–24. https://doi.org/10.1109/ISSSR65654.2025.00062.
  11. Poole, Nick. 2015. ‘Guest Blog: Aggregation & the Culture Grid’. Museums Computer Group, May 31. https://museumscomputergroup.org.uk/culture-grid/.
  12. Rawson, Katie, and Trevor Muñoz. 2016. ‘Against Cleaning’. Curating Menus, July 6. http://www.curatingmenus.org/articles/against-cleaning/.
  13. ‘Spectrum’. n.d. Collections Trust. Accessed 7 May 2026. https://collectionstrust.org.uk/spectrum/.
  14. Towards a National Collection. 2025. ‘N-RICH Prototype’. July. https://www.nationalcollection.org.uk/n-rich-prototype.