Daejeon, July 27–31
Abstract. The notion of democracy and the functions of the state of the law in modern liberal countries are based on legislation and agreements that are discussed and decided in democratically elected national parliaments. In an analogous way, various international organizations have been established to discuss and agree upon international matters, such as state of the law principles, war and peace, climate change, global health issues, and food and agriculture. Examples of such international organizations include the Parliament of EU, United Nations (UN), Inter-Parliamentary Union (IPU), World Health Organization (WHO), Red Cross, and Food and Agriculture Organization (FAO) to name a few. This paper presents a Linked Open Data (LOD) approach for publishing and studying assembly minutes data of international organizations, based on lessons learned in developing analogous systems for parliamentary discussion minutes. As a case study, minutes of the Assembly of the League of Nations (LoN) (1920–1946) are considered. A new system League of Nations Sampo is presented, re-using the ParliamentSampo framework for the speeches and prosopography of the Parliament of Finland (1907–). League of Nations Sampo is based 27,000 pages of minutes of LoN assembly meetings, a prosopographical knowledge graph about some 3,100 people mentioned in the minutes, and on contextualizing data about the real world.
Keywords: knowledge graphs, digital humanities, information retrieval, data analysis
To make decision making in parliaments and international organizations transparent, minutes of discussions among the decision makers are documented and made open to the public and researchers as printed books and as PDF and other documents online. This paper argues that opening the documents for people to read is not enough. Even if the documents are openly available, they are typically not Findable, Accessible, lnteroperable, and Re-usable according to the FAIR principles (Wilkinson et al., 2016). As a remedy, the minutes should be published not only in human-readable form but also as semantic data for computers to interpret and use, including contextual data about the decision makers and organizations involved, as well as about the underlying real world. This would facilitate development of easy-to-use intelligent applications for scholarly Digital Humanities (DH) research and public engagement.
From a DH point of view, minutes data differ from organization to organization. ln national parliaments, the minutes typically record the speeches of the parliamentarians word by word. For example, the Parliament of Finland corpus (Hyvönen et al., 2025) contains textual transcripts of over million speeches since 1907. On the other hand, in the assembly minutes of the League of Nations (1920–1946) (LONTAD, 2022; Leskinen et al., 2026), the focus is more on documenting the decisions. However, in all cases the minutes documents are essentially texts making references to various topical matters, people, organizations, and places in time. Our research hypothesis is therefore that a similar kind of technical solution can be applied to different kind of minutes cases for FAIRifying the data and for building applications on top of the data.
Parliaments have created speech corpora and datasets of both historical and contemporary parliaments. Parliamentary data have been used in many fields of research, such as linguistics, political science, legal studies, media studies, economics, and history, including long-term studies (Guldi, 2019). Semantic web technologies have been applied for linking and enriching parliamentary data (Van Aggelen et al., 2017). In the same way, international organizations have produced vast corpora of assembly minutes, plenary debates, committee reports, and voting records for accountability and global governance (Cafiero et al., 2025). The vision in our work is to turn such dispersed materials into a sustainable LOD ecosystem (Cafiero, 2023; Hyvönen et al., 2025) that allows users to trace who spoke, on what issues, in which capacity, and how agendas evolved over time, while remaining explicit about uncertainties, biases, and curatorial choices. Our goal is to create and align with each other minutes data from several related international organizations, with a focus on Geneva-based ones (MM Project, 2026).
To test and demonstrate these ideas, we present the case of the League of Nations (LoN) based on data from the LONTAD project (LONTAD, 2022; Wells, 2019) enriched from other related data sources. These materials have been accessible online for humans to read, but not as data for research and application development.
We transformed the 27,000 pages of LoN Assembly minutes related to some 3,100 mentioned representatives and other people into a LOD service with a SPARQL endpoint (LoN LOD Service, 2026) and a semantic portal League of Nations Sampo (LoN Portal, 2026). Proposographical data about the actors and the real world (cf. Figure 1) were extracted and aligned from the related data services of Figure 2.
The League of Nations Sampo SPARQL endpoint can be used directly by scripting in DH research and for application development. To demonstrate this, League of Nations Sampo portal was built by using the Sampo model (Hyvönen, 2023) and Sampo-Ul framework (Ikkala et al., 2021; Rantala et al., 2023) and the domain specific ParliamentSampo framework (Hyvönen et el., 2025).
The idea of the "Sampo framework" is to take an existing Sampo in a domain of interest, with a ready to use user interface (UI) model and knowledge graph as a starting point and modify it instead of starting from scratch. This approach allows for extremely rapid software development, if the framework used is mostly fit for the new purpose (Sampo-UI, 2026). See (Leskinen et al., 2026) for more details about the League of Nations Sampo LOD service and portal application available openly on the Web. According to first informal evaluations of domain experts, the data quality and the portal seem useful. Usability of the Sampo-Ul model from an end user viewpoint has been evaluated in a related Sampo project (Burrows et al., 2020).
Acknowledgments. We thank the LONTAD archives project, UN Library & Archives Geneva, Lonsea project, Dodis, and Metagrid for fruitful collaborations. Support of the Finnish DH research infrastructure initiative FIN-CLARIAH/DARIAH-FI funded by the Research Council of Finland and NextConnectionEU is acknowledged. Computational resources of CSC IT Center for Science of Finland have been used.