Daejeon, July 27–31
The recent mass-scraping of digital content, and the resulting backlash from researchers whose work has been used to train AI tools, has “led to disquiet among both open-access advocates and detractors” (Eve 2025). Open license initiatives have seemingly given “our collections wholesale to AI as part of mass training… inadvertently” (Cohen 2025). At the same time, cultural heritage institutions and digital scholarship providers are seeing their platforms crippled by unethical, predatory scraping by generative AI providers (Weinberg 2025). This creates a new tension: many in the Digital Humanities wish to abide by the FAIR data principles in order to make both code and data discoverable, accessible, interoperable, and reusable (Wilkinson et al 2016) in the pursuit of “open science” (Vicente-Saez and Martinez-Fuentes 2018). However, many are now adopting the approach of being “as open as possible, as closed as necessary”. This phrase originally appears in the FAIR data principles to discuss the challenges of openly licensing data abiding by GDPR and privacy legislation (Landi et. al. 2020) but is now being reused within the library and digital humanities sector (Bryant 2025) to describe the boundaries that need to be put in place to ensure longevity of infrastructure, and protection of digital collections and data quality, in a time of predatory generative AI.
The poster describes how one scholarly infrastructure, Transkribus (an Artificial Intelligence platform for Automated Text Recognition and information extraction from historical documents ( https://www.transkribus.org, Romein et.al. 2025)) is navigating these opposing forces between openness and protection. Transkribus is maintained and developed by READ-COOP ( https://readcoop.org), established in 2019 as a European Cooperative Society. Since its inception, the cooperative has navigated a critical tension for responsible digital cultural heritage infrastructure providers: balancing community imperatives of development, openness, and accessibility with the practical necessities of data protection, intellectual property preservation, and financial sustainability. The business model requires selective closure of the Transkribus infrastructure, in order to protect the proprietary algorithms and extensive training datasets that form READ-COOP’s core intellectual property: assets developed over years with significant public investment. Without this protection, larger technology providers could re-appropriate these resources, threatening the cooperative’s viability and its ability to serve its community. This has attracted criticism from those researchers wedded to open science approaches, although few could use the large (400TB) models at the heart of Transkribus, if they were openly licensed, given the computational resources and skillset required. Objections seem based on principle, rather than practicalities.
However, READ-COOP does remain committed to maximum accessibility through multiple mechanisms. Users retain full ownership of their data and can export all materials in open formats. The cooperative has published 300 licensed AI models for internal community use and encourages members to share training data through repositories such as HTR-United. READ-COOP also offers free monthly credits to enable trial use and support hobbyist researchers. A comprehensive Scholarship Programme provides free access to students from 73 countries, having supported 336 individuals with over one million free credits by October 2024 (Nockels, Gooding, and Terras 2025). The cooperative structure itself embodies openness through democratic governance, enabling members to provide meaningful input into platform development and strategic decisions (Terras et. al. 2025). Monthly meetings, active communication channels, and annual conferences foster transparent dialogue between users and developers. This approach demonstrates that ethical AI governance requires not merely technical openness but also structural mechanisms that ensure community participation, accountability, and shared ownership, thereby safeguarding the systems’ future. Transkribus may offer a walled garden, but it is the Digital Humanities’ walled garden.
READ-COOP’s experience suggests that responsible platform development now necessitates pragmatic balance rather than an absolutist position to openness in the digital sphere. By prioritising community needs, transparent governance, and sustainable business practices over profit maximisation, the cooperative model offers an alternative framework for AI infrastructure—one that protects both the technology and the community it serves, demonstrating that closure and openness need not be oppositional but can be strategically calibrated to ensure long-term viability and broad accessibility. We suggest that this may be the reality of openness in the current digital age: reframing our approach to ensure our digital assets are not stripped from us by unethical data scraping. This requires us to re-calibrate and reframe scholarly expectations, and to seek alternative approaches by which we can build and own data infrastructures ourselves.