DH 2026

Daejeon, July 27–31

Thu, July 3009:00–10:30S055209-211
Short Paper

From paper to cloud: Embracing archival heritage through Digital Humanities

Magally Alegre Henderson
Pontificia Universidad Católica del Perú, Peru · malegreh@pucp.pe

Introduction

The Pontifical Catholic University of Peru (PUCP) holds the country's most valuable private rare books library and historical archive, only comparable to the national repositories. Only a fraction of these extensive collections has been digitized and are held in the PUCP Institutional Repository and the Open Data Portal. The minimal digital accessibility to this rich documentary heritage, the absence of institutional digitization and cataloguing standards, and the lack of textual recognition protocols in the digitization of archival and historical sources in Peru, printed or handwritten, pose a serious limitation for students, professors, and researchers, at the PUCP and beyond, and challenges the development of Digital Humanities in Peruvian Studies.

A digital gap

A situation shared by much of the GLAM sector in Peru, this digital gap, both in private and public collections, is part of a broader pattern of neglect and limited access to the country's documentary patrimony. For example, the General National Archive in Peru has been at risk of eviction from its historic premises over the last two years. It also explains why Peruvian archives have received almost twice as many grants as any other country in the region for rescuing documentary collections through the British Library Endangered Archives Programme (EAP), which with the support of Arcadia promotes the digitization of endangered documentary heritage. Most importantly, Peru lacks a national search engine that integrates the catalogs and digital objects of state-held heritage collections, nor does it have a national cataloguing and digitization policy aimed at ensuring interoperability of this data. For example, the AtoM (Access to Memory) system implemented in 2024 at the Peruvian National Archives (https://fondosdocumentales.agn.gob.pe/) houses a limited portion of historical records digitized with different protocols in the production of images and the generation of metadata. Some documents are published in a low resolution which compromises legilibility. In this regard, we propose that, through Digital Humanities methodologies, university archives can help address this issue by ensuring digital accessibility to their collections and guaranteeing the right to the cultural heritage they preserve.

Documentary heritage engagement

In line with this mission, the Historical Archive of the Riva-Agüero Institute (AHRA) at PUCP is firmly committed to ensuring full digital accessibility, so that our valuable documentary heritage is available to the academic community both in Peru and abroad. We are actively engaged in Digital Humanities initiatives based on three fundamental pillars: (1) training, (2) project design, and (3) international partnerships, the results of which I will briefly summarize below.

The Summer Courses in Digital Humanities PREFALC CONTAMAZ - Contar la Amazonía (Telling the Amazon) organized last year jointly with Marie and Louis Pasteur University (France) and the National University of Colombia evidences our commitment to training. Thanks to the sponsorship of the France Latin America-Caribbean Regional Program, students and scholars from PUCP and UNAL experienced how to use digital humanities tools to analyze an unpublished source, a set of 24 notebooks by Marcos Cavero, an administrator at the Iquitos River Customs Office, handwritten in the early 20th century. The notebooks, which belong to the AHRA Félix Denegri Luna collection, have been digitized thanks to the EAP1495 Peruvian Legacy project, funded by the British Library’s Endangered Archives Programme. A second edition was dedicated to the digital processing of José Eustasio Rivera’s La Vorágine. Both sets of manuscripts were used to teach digital processing, including data extraction with Transkribus, data analysis with Voyant Tools, Iramuteq, and uMap, and other developments in TEI/XML text editing.

The AI-2 Archiving project at the AHRA is developing an Artificial Intelligence toolkit for the standardization and cleaning of its databases, accelerating the process of implementing an open-source cataloguing software (AtoM) for the registration of documentary collections in the AHRA. The diversity of media and types of documents housed in the 86 collections of our archive, whose oldest documents date back to the 16th century, makes it an exceptional pilot for the development of a cataloguing system that can be used in other university documentary and heritage collections, with the expectation that it can be implemented in all types of Peruvian heritage collections.

Ai2 Archiving is in the design phase of an AI-based software prototype to create a customized cataloging system tailored to the specific needs of repositories in Peru’s GLAM sector (Galleries, Libraries, Archives, and Museums). This prototype consists of a toolkit that includes manuals, prompts, and training materials in Spanish, focused on the cataloging of documentary heritage. Ai2 Archiving improves access to historical archives by speeding up archival cataloging processes, which are key to the preservation and accessibility of historical and documentary collections but are typically laborious and time-consuming. This challenge is even greater in the case of historical manuscripts, which require additional processing time for manual paleographic transcription or the use of Handwritten Text Recognition (HTR) technologies.

In summary

To address the challenges of standardization—which guarantees data interoperability—in the cataloging of documentary collections and other cultural heritage repositories, it is essential to carry out manual data cleaning and categorization processes that require a considerable number of hours and intensive effort. Although artificial intelligence has begun to be used in cataloging, accelerating the processing of repetitive tasks and the organization of large volumes of information, its application to specific cataloging workflows in Peruvian archives is still in its infancy. In the absence of thesaurus or ontologies specifically designed for cataloging Peruvian cultural heritage, this gap underscores the urgent need to develop and adapt AI technologies that not only understand but also efficiently support data extraction and systematization, thereby enhancing the digital accessibility and preservation of Peru’s archival heritage.

As the Latin American hub for the British Library’s Endangered Archives Programme, AHRA has also gained significant regional recognition for promoting digital humanities. Our commitment to advancing accessibility to documentary heritage within university archives addresses the challenges of accessing digital sources and the interoperability needs of digital repositories for the study of Peru’s cultural legacy.

References
  1. Alegre Henderson, Magally / Montoya Blua, Verónica / Sifuentes Arroyo, Raúl / Vera Zúñiga, Javier / Luque Balbuena, Sandra / Castro Aguilar, José / Grijalba Touzett, Rodrigo (2025): “AI-2 Archiving: Validation of a specialized toolkit for cataloguing historical archives using artificial intelligence.” Concurso Anual de Proyectos CAP Innovation 2024 Pi1150. Vice Presidency for Research. Pontificia Universidad Católica del Perú.
  2. Alves, Daniel (2014): “Guest Editor's Introduction: Digital Methods and Tools for Historical Research,” International Journal of Humanities and Arts Computing 8(1): 1-12. <https://doi.org/10.3366/ijhac.2014.0116>.
  3. Arano, Silvia (2005): “Los tesauros y las ontologías en la Biblioteconomía y la Documentación”, Hipertext.net, 3 <https://arxiu-web.upf.edu/hipertextnet/numero-3/tesauros.html>.
  4. Bushey, Jessica (2012) “ICA-AtoM: open-source software for archival description.” Archivi & Computer 22, 1: 10-25.
  5. Colavizza, Giovanni, et al. (2021): “Archives and AI: An Overview of Current Debates and Future Perspectives,” ACM Journal on Computing and Cultural Heritage 15(1): 15 pp. <https://doi.org/10.1145/3479010>.
  6. Dahl, Christian M, et al. (2023): “Applications of machine learning in tabular document digitisation,” Historical Methods: A Journal of Quantitative and Interdisciplinary History 56(1): 34-48 <https://doi-org/10.1080/01615440.2023.2164879>.
  7. Eide, Ø. “Modelling and networks in digital humanities.” In Routledge eBooks, 2020, pp. 91–108. <https://doi.org/10.4324/9780429777028-8>.
  8. Erhard, Franz X. (2025): “Text and Layout Recognition for Tibetan Newspapers with Transkribus,” Revue d’Etudes Tibétaines 74: 128-171.
  9. Flor de Sousa Neto, Arthur, et al. (2024): “Data Augmentation for Offline Handwritten Text Recognition: A Systematic Literature Review,” SN Computer Science 5(258): 1-20. <https://doi.org/10.1007/s42979-023-02583-6>.
  10. García Serrano, Ana / Menta Garuz, Antonio (2022): “La inteligencia artificial en las Humanidades Digitales: dos experiencias con corpus digitales,” Revista de Humanidades Digitales 7: 19-39 <https://doi.org/10.5944/rhd.vol.7.2022.30928>.
  11. Guerrero, Dalia (2016): “Humanidades Digitais: novos desafíos e oportunidades,” Revista Internacional del Libro, Digitalización y Bibliotecas 2(2): 16 pp. <https://doi.org/10.37467/gkarevdig.v2.779>.
  12. Jaillant, Lise (2022): “How can we make born‑digital and digitised archives more accessible? Identifying obstacles and solutions,” Archival Science 22: 417-436 <https://doi.org/10.1007/s10502-022-09390-7>.
  13. Jiménez Bolívar, Mercedes (2020): “El modelo conceptual de la descripción archivística y el proyecto ATOM (acceso a la memoria) Un ejemplo, ATOM en el archivo fotográfico de la Universidad de Málaga.” Revista TRIA 24: 67-98 <https://dialnet.unirioja.es/servlet/articulo?codigo=8262683>.
  14. Markus Müller (2021, 23. Juli). Uncovering censorship in the 16th century with Transkribus and Python. Episode II: Let Python speak to Transkribus. Digital Humanities Lab. Abgerufen am 5. April 2024, <https://doi.org/10.58079/nl5y>.
  15. Marsili, Giulia / Hassam, Stephan N. (2025): “From archives to museum and back: transcribing, digitising and enriching cultural heritage and manuscript legacy data of the Villa del Casale of Piazza Armerina,” Digital Applications in Archaeology and Cultural Heritage 38: 1-15 <https://doi.org/10.1016/j.daach.2025.e00441>.
  16. Naumis Peña, Catalina (2007): Estudio comparativo de tesauros bibliotecológicos en lengua española, Investigación Bibliotecológica 21, 42: 195-210, <http://www.scielo.org.mx/scielo.php?script=sci_arttext&pid=S0187-358X2007000100009&lng=es&tlng=es>
  17. Nockels, Joe / Gooding, Paul / Ames, Sarah / Terras, Melissa (2022): “Understanding the application of handwritten text recognition technology in heritage contexts: a systematic review of Transkribus in published research”, Archival Science 22: 367–392. <https://doi.org/10.1007/s10502-022-09397-0>.
  18. Parra Mora, Gloria María / Hilda Paola Serna González (2020): “La descripción documental como medio de acceso a la memoria institucional: experiencia de uso del software ICA AtoM en la Fundación Universitaria para el Desarrollo Humano –UNINPAHU- Bogotá, Colombia.” e-Ciencias de la Información <https://doi.org/10.15517/eci.v10i1.39774>.
  19. Peña Pimentel, Miriam (2020): “Lecturas y contenidos: las bibliotecas digitales,” Bibliographica 3(2): 187-206 <https://doi.org/10.22201/iib.2594178xe.2020.2.82>.
  20. Pippi, Vittorio. et al. (2023): “How to Choose Pretrained Handwriting Recognition Models for Single Writer Fine-Tuning,” in Fink, Gernot A. et al., Document Analysis and Recognition – ICDAR 2023 Part II <https://doi.org/10.1007/978-3-031-41679-8>.
  21. Raventós Pajares, Pepita, et al. (2022): “AI and archive. Handwritten Text Recognition Applied to Patrimonial Holdings: An Example of 10 diaries written by Spanish Republican Teachers in 1932,” IEEE International Conference on Big Data, 10: 2572-2577. <https://doi.org/10.14296/1220.9781912250356>.
  22. Riva, Pat/ Patrick Le Bœuf/ Maja Žumer (2017):IFLA Library Reference Model. A Conceptual Model for Bibliographic Information,” International Federation of Library Associations and Institutions, <https://www.ifla.org/wp-content/uploads/2019/05/assets/cataloguing/frbr-lrm/ifla-lrm- august-2017_rev201712.pdf>.
  23. Rojas Castro, Antonio (2020): “FAIR enough? Building DH Resources in an Unequal World.” Presentation at Digital Humanities-Kolloquium, Berlin <http://dx.doi.org/10.17613/1xxn-gj36>; Consortium of European Social Science Data Archives, Data Archiving Guide, <https://dag.cessda.eu/>.
  24. Romein, C. A, et al. (2024b): “From research proposal to project management. A guide from the Transkribus community on planning and executing workflows for researchers and GLAM-professionals.” International Journal of Digital Humanities 7(2), published online 01.09.2025 <https://doi.org/10.1007/s42803-025-00107-7>.
  25. Romein, C. A., et al. (2024a): “Assessing advanced handwritten text recognition engines for digitizing historical documents.” International Journal of Digital Humanities, 7(1): 115-134. <https://doi.org/10.1007/s42803-025-00100-0>.
  26. Salamandra Palacios, Wesly (2022): “Mejoras para la implementación del software AtoM (Access to Memory) en un archivo.” Master’s Thesis in Information Management, Valencia: Universitat Politècnica de València <http://hdl.handle.net/10251/190618>.
  27. Shadya Sanchez Carrera/ Roberto Zariquiey / Arturo Oncevay. “Unlocking Knowledge with OCR-Driven Document Digitization for Peruvian Indigenous Languages,” in Proceedings of the 4th Workshop on Natural Language Processing for Indigenous Languages of the Americas (Americas NLP 2024): 103–111, CDMX: Association for Computational Linguistics, 2024.
  28. Suárez, Susana/ Rocío Aguilera (2019): “Del Excel al ICA-Atom” in Imágenes de una escuela con archivo histórico: escuela cooperativa Amuyén: 29–37. Universidad Nacional de Mar del Plata <http://humadoc.mdp.edu.ar:8080/handle/123456789/854>.
  29. UNESCO (2016): Recommendation concerning the Preservation of, and Access to, Documentary Heritage Including in Digital Form.
  30. Van Garderen, Peter (2009): The ICA-AtoM Project and Technology. Archives Association of Brazil presentation <https://wiki.accesstomemory.org/wiki/File:VanGarderen-ICA-AtoM-2009.pdf>.