Daejeon, July 27–31
Interactions with simulated historical figures often lead to the uncritical acceptance of their narratives. This creates the false impression that historical events are more settled than they actually were. Such outcomes are evident in initiatives such as Digital Einstein and MIT's Living Memories, both interactive digital projects that allow users to engage with representations of historical people or experiences. The absence of critical engagement highlights a deeper issue: historians' interpretive work differs from the outputs of machines connected to knowledge bases. Historians rigorously interpret fragmentary and often contradictory evidence, a process that requires careful examination and weighing of partial and conflicting information. In contrast, this interpretive process is entirely concealed within AI systems. As a result, users may mistakenly perceive digital outputs as historical truth. This phenomenon is called "digital positivism," meaning the tendency to treat digital artifacts, such as AI-generated responses or simulations, as being as reliable or authentic as original historical sources.
To address this challenge, it is necessary to determine whether actionable steps exist and, if so, to precisely identify them. This question leads directly into the context of the present case study, which seeks to offer a concrete response.
This case study focuses on King Sejong the Great, the renowned ruler of Joseon Korea (1397-1450). He commissioned the creation of Hangul, the Korean alphabet. His reign is extensively documented in the Sillok, or court records, which are detailed annals kept by the royal court. These records provide a robust foundation for evaluating epistemological design, meaning the structure for how historical knowledge is formed and justified. This paper introduces a four-tier epistemic architecture, a framework for categorizing different types of historical knowledge. It compels a Generative AI model (an artificial intelligence system that creates content) to distinguish among four categories of historical knowledge: verified facts, widely held interpretations, plausible inferences, and genuinely contested debates. The system incorporates a retrieval-augmented generation (RAG) subsystem that enables the AI to access external sources during its generation process. This ensures that distinctions between historical certainty and historiographical reasoning remain explicit.
This framework makes three main contributions. First, it shows that computational methods, automated processes, or algorithms used for data analysis and decision-making can preserve interpretive complexity, thereby maintaining nuanced, multi-layered understandings in AI-driven historical dialogue. Second, it introduces a decolonial governance model, granting Korean cultural institutions substantive authority and decision-making power over heritage representation, moving beyond a merely consultative role. Third, it provides methodological guidance on epistemic responsibility, the obligation to act knowledgeably and ethically, in AI systems, and offers practical recommendations for digital humanities practitioners facing similar challenges. The central argument is that digital humanities can rigorously employ AI without compromising historical integrity or cultural accountability, provided that deliberate architectural decisions are made, an approach uncommon in current systems.
Keywords: conversational AI, digital humanities, historical uncertainty, King Sejong the Great, retrieval-augmented generation, decolonial AI, Korean cultural heritage
Can AI recreate the voice of historical figures? This question goes to the heart of what digital humanities scholars worry about: whether computers can help us understand the past without oversimplifications. Recent projects show both the promise and the problems. Digital Einstein answered over 100,000 questions in two weeks. MIT's Living Memories created AI versions of Leonardo da Vinci and Murasaki Shikibu to help students learn. However, these systems all share a common flaw: they make history appear more certain than it actually is.
The problem is fundamental. When AI generates text that sounds like a historical figure speaking, users tend to forget they are interacting with a computer's interpretation, not the person's actual words. This is especially problematic because the AI sounds so confident. As researcher Ted Underwood recently showed, even AI models trained specifically on historical texts produce subtle anachronisms, using modern concepts and phrasings that didn't exist in the past.
This tension defines modern digital humanities work. Alan Liu argues that when we use computational methods, we must preserve the complexity and ambiguity that make the humanities valuable. Roopika Risam warns that digital projects can repeat colonial patterns when Western institutions control how non-Western cultures are represented. These aren't just theoretical concerns; they affect how we design and build AI systems.
We could build a generic AI tutor that explains Joseon Dynasty history in a neutral manner. However, recent research on immersive learning and narrative pedagogy demonstrates that culturally grounded personas offer distinct cognitive and affective advantages that generic tutors cannot replicate. Studies of XR docent systems and museum chatbots reveal that personas create an embodied historical presence, enabling visitors to form emotional connections that enhance memory retention and foster deeper engagement with cultural contexts.
With this understanding, a King Sejong persona specifically offers unique educational affordances:
While these are clear pedagogical advantages, they come with responsibility. A persona that sounds authentic risks misleading users about the certainty of historical knowledge. This is precisely what prompted us to create the epistemic framework presented in this paper, to preserve the engagement benefits of persona-based learning while maintaining scholarly integrity regarding what we can and cannot know.
In this work, we focus on King Sejong the Great (세종대왕, 1397-1450), who ruled the Joseon Dynasty of Korea for 32 years and created Hangul, the Korean alphabet. Sejong is an ideal case for several reasons:
Rather than building yet another chatbot, we developed an epistemic framework, a system for representing what we know and how certain we are. The objective of our four-tier system is to discriminate between:
This approach builds on the work of Michael Piotrowski, who argues that digital humanities must find ways to formally represent uncertainty rather than hide it.
Several recent projects demonstrate both the possibilities and challenges of recreating historical figures with AI.
UneeQ Digital Einstein (2021) combined a computer-generated face, a synthesized voice, and integration with scientific computation tools. The project generated significant traffic, with over 100,000 interactions in just two weeks. But the conversations stayed shallow. Einstein could answer questions about his documented life and known scientific work, but he was unable to discuss anything after his death in 1955 or engage emotionally in meaningful ways. While the project succeeded as a marketing tool, it failed as a means of historical education.
MIT's Living Memories (2023) took a more scholarly approach. Researchers created AI versions of several historical figures and tested them with 90 students. Students learned more and felt more motivated compared to traditional methods. However, the researchers identified persistent problems: the AI confidently fabricated plausible-sounding facts, filled evidence gaps with fiction, and struggled to strike a balance between engagement and accuracy. Technical sophistication alone wasn't enough; researchers needed robust fact-checking and clear warnings about AI limitations.
The Mary Sibley Project (2024) at Lindenwood University offers the most relevant lessons. Researchers trained ChatGPT on Sibley's complete diaries from 19th-century Missouri. The AI successfully captured her religious tone and old-fashioned language. However, success required university archivists to review every response; they caught the AI confidently stating incorrect founding dates and institutional names. The project also added that Sibley's historically accurate views on race and gender would trouble modern readers. This points to a more mature approach than either hiding an uncomfortable history or presenting it without context, although the tensions remain unresolved.
Most relevant to our case, the Korean Institute of Science and Technology deployed an AI system for Joseon Dynasty historical records in 2025. The system provides answers to questions about 472 years of Korean history, accompanied by clear source citations. This proves that AI can handle Korean historical materials. However, it also revealed ongoing challenges, including managing dense chronological texts, balancing facts with interpretation, and processing massive datasets (over 336,000 career records with annotations).
When polled together, here are the learnings we can derive from these projects:
What works:
What consistently fails:
As Underwood's recent research confirms, perfect historical authenticity may be technically impossible. Even AI models trained exclusively on historical texts produce detectable anachronisms. This means we need realistic expectations: our systems create interpretations based on historical sources, not authentic recreations of historical consciousness.
DH scholars have long warned about these problems. Johanna Drucker argues that we shouldn't treat digitized materials as transparent windows onto the past. Every digital representation involves interpretive choices, such as which documents to include, how to organize them, and what metadata to add. These choices shape what users can find and understand.
Roopika Risam extends this critique to show how digital projects can perpetuate colonial power structures. When Western institutions digitize non-Western cultural materials, benefits often flow outward. Western researchers gain access, prestige, and publications, while source communities see little return. Even "universal" technical standards often encode Western assumptions about how knowledge should be organized.
For AI specifically, we must design systems that preserve complexity rather than eliminate it. The goal isn't to make history simpler or more accessible at the cost of accuracy; it is to make complexity navigable.
Our primary contribution is a framework that enables AI to distinguish among types of historical knowledge and to address the core problem: current AI systems treat all generated text as equally valid, whether it states documented facts or makes educated guesses.
At the most basic level, we consider historical knowledge to exist in layers, from the most certain to the most debatable. From this perspective, we can identify the following tiers:
TIER 1: DOCUMENTED FACTS (HIGH confidence)
These represent statements that we can consider factual, as they are verifiable directly in primary sources and supported by multiple independent sources.
Example : "When were you born?"
AI Response : "I was born in 1397, the 26th year of King Taejo's reign."
What the user sees:
Sources : 3 primary sources confirm this.
Confidence : HIGH (documented fact)
Citation : Sejong Sillok, Biographical Section; Veritable Records of King Taejo, Year 26
This is straightforward; the birth year appears in multiple historical records written during or shortly after Sejong's life. However, in the background, the system searches its knowledge base for excerpts from documents mentioning Sejong's birth. It identifies the relevant passages, verifies their consistency, and generates a response that utilizes the language style of the historical sources. The metadata tags indicate that this is a biographical fact with high certainty.
TIER 2: SCHOLARLY CONSENSUS (MEDIUM-HIGH confidence)
These are interpretations that require inference but are widely accepted by historians. The evidence is strong, even though it requires some interpretation.
Example : "Why did you create Hangul?"
AI Response : " The spoken language of our country is different from that of China and does not suit the Chinese characters. Therefore, amongst uneducated people, many have been unable to express their feelings in writing when they have something they wish to put into words. I was greatly distressed by this, and so I have made twenty-eight new letters to everyone to practice them at their ease, and adapt them to their daily lives. "
What the user sees:
Sources : Hunminjeongeum preface (primary source) + 4 scholarly interpretations
Confidence : MEDIUM-HIGH (most historians agree)
Citation : Hunminjeongeum (1446), Preface by King Sejong
Note : We infer Sejong's motivation from his explicit statements in the preface of his work and from his literacy policies. Most scholars agree on the democratic/practical motivations, though some emphasize additional factors, such as linguistic nationalism.
In this case, the system retrieved the Hunminjeongeum preface (a primary source) along with several scholarly articles about the creation of Hangul. It examines the degree of consensus among historians regarding the interpretation. Because most of the retrieved articles support the democratic motivation explanation, the system classifies this as a scholarly consensus.
TIER 3: REASONABLE SPECULATION (MEDIUM-LOW confidence)
Answers belong to this tier when they are considered educated guesses based on historical context, but with limited direct evidence. Historians consider them plausible but can't prove them definitively.
Example : "How did you feel when yangban scholars opposed Hangul?"
A I Response : "While I cannot directly describe my feelings, court records document my decisions and policies, not inner emotional states; historical context suggests this was challenging. The records show I proceeded with Hangul despite elite opposition and assigned Hall of Worthies scholars to develop theoretical justifications. This determination suggests I viewed Hangul as essential to good governance even when facing resistance from officials who saw Chinese characters as culturally superior."
What the user sees:
Sources : Inferred from court records + historical context
Confidence : MEDIUM-LOW (reasonable speculation)
Citation : Sejong Sillok, Years 25-28 (Hangul development period
Note : Sejong's emotional states are rarely directly documented in court records. We infer from documented actions (proceeding despite opposition) and policy patterns. Historians consider this plausible but acknowledge limited direct evidence. In the historical context in which these documents were collected, we find the Joseon court records prioritized policy documentation over personal emotion, following Neo-Confucian traditions.
Notice the AI explicitly states it cannot know Sejong's feelings. Instead, it explains what evidence exists (documented actions) and why historians make this inference. This model's good historical thinking involves distinguishing between what we know and what we reasonably infer.
The system classifies the query as an emotion/motivation inquiry and searches for direct statements about Sejong's feelings, and finds none. It retrieves documents about his actions during this period and uses those to construct a contextualized inference. The response explicitly flags its speculative status.
TIER 4: CONTESTED INTERPRETATIONS (Multiple views, active debate)
It is inevitable that historians will disagree on certain questions, with different scholars using different evidence or frameworks to reach different conclusions.
Example : "Did you create Hangul alone or with your scholars?"
AI Response : "Historians actively debate this question. The evidence supports multiple interpretations.”
This response was guided by various viewpoints identified from the context of this question:
Majority View (Collaborative Creation, ~65% of scholars): King Sejong worked closely with Hall of Worthies scholars, especially Prince Suyang and scholars Jeong In-ji, Choe Hang, and Seong Sam-mun. The Hunminjeongeum Haerye (Explanatory Guide) credits these scholars with theoretical elaboration. This view emphasizes Confucian governance ideals, where rulers work in conjunction with learned advisors.
Minority View (Sole Creator, ~35% of scholars): King Sejong personally designed Hangul's phonetic system, with scholars only writing explanatory materials afterward. Evidence includes the preface's first-person statements ('I have created 28 letters') and Hangul's linguistic sophistication, suggesting a single brilliant mind rather than a committee design.
Both interpretations use the same sources but emphasize different aspects. The collaborative view stresses governance patterns and scholar credits. The sole-creator view emphasizes linguistic sophistication and first-person language, resulting in an interpretation that depends on whether one prioritizes social-historical context or technical linguistic analysis.
What the user sees:
Sources : Same primary sources interpreted differently; 8 scholarly works divided
Confidence : CONTESTED (active debate, roughly 65-35% split)
Citations : Hunminjeongeum (1446), Lee & Ramsey (2011), containing a collaborative argument, and Kim-Renaud (1997), pointing to a sole creator argument
Note : This isn't resolvable with better AI or more data; it reflects genuine historiographical disagreement. Different scholars employ different frameworks (social history versus linguistic analysis), resulting in varying conclusions drawn from the same evidence.
We can consider this example as the most sophisticated response type. The AI doesn't pick a side or try to synthesize a middle position. Instead, it presents both views fairly and explains why historians disagree. This teaches users that history involves interpretation, not just fact-gathering.
During the search for scholarly literature, the AI detects significant disagreement (not 90-10%, but closer to 65-35%). It retrieves representative articles from each camp and analyzes their reasoning. The response explains both positions and the methodological differences that lead to different conclusions.
Figure 1 : System workflow from question to response, including ongoing review by Korean historians.
While most AI chatbots implement the classic RAG workflow in which the AI generates a response based on the training or on a Retrieval Augmented Generation (RAG) architecture, the proposed solution adds a classification sub-system to identify the certainty level:
These additional steps are at the core of our contribution:
This approach builds on Johanna Drucker's argument that digital humanities should make interpretive layers visible, rather than hiding them behind technical sophistication.
King Sejong presents an ideal case for developing and testing our framework for several interconnected reasons.
Exceptional historical documentation
The Veritable Records of the Joseon Dynasty document Sejong's 32-year reign in extraordinary detail, comprising 888 books that cover daily court activities, policy decisions, and scholarly discussions. Eight official historians accompanied the king, providing legal protection to record even the most embarrassing incidents. This comprehensive coverage was recognized by UNESCO as "Memory of the World" in 1997. Unlike many historical figures, whose lives only provide fragmented evidence at best, Sejong's life is extensively documented.
Already digitized and accessible
These records have been fully digitized since 2006 and are freely available online at sillok.history.go.kr, in both the original Classical Chinese and a modern Korean translation. English translation is underway (expected completion 2033). The texts employ standardized TEI-compliant XML encoding, which facilitates computational processing. The Korean Institute of Science and Technology already deployed an AI system for these records in 2025, proving technical feasibility.
Active scholarly debates
Historians continue to debate key questions about Sejong: Did he invent Hangul alone or in collaboration with others? What linguistic influences shaped Hangul's design? What motivated his literacy policies? These live debates let us test how well Tier 4 (contested interpretations) works. We can demonstrate the system handling genuine historiographical disagreement, not manufactured examples.
Cultural significance requires community partnership
Sejong remains profoundly important to Korean national identity. Hangul Day (October 9) is a national public holiday. His Birth Commemoration Day was officially designated in 2024. His statue stands prominently at Gwanghwamun Square in the center of Seoul. He appears on the 10,000 won banknote. There are 248 King Sejong Institutes in 85 countries, teaching Korean worldwide. A 2024 Gallup Korea survey ranked Sejong second among the most respected historical figures. This deep cultural investment means any AI representation must involve Korean communities as decision-makers with veto power, not just consultants.
Table 1 shows how different types of Sejong sources map to our tier system.
Cultural context requiring ethical care
Korean ancestor veneration traditions (jesa, charye rituals) involve belief in a spiritual connection with deceased ancestors and specific protocols for respectfully "calling" ancestral souls. From this cultural perspective, AI simulation may be inappropriate—not just poorly implemented, but fundamentally flawed, regardless of its technical quality. This means that Korean communities must determine whether any AI representation aligns with proper ancestor worship and memorialization. External researchers cannot make this decision.
While full implementation requires sustained participatory research, we establish here a three-tier governance structure making the Korean authority operational:
We propose a phased participatory research approach, beginning with a cultural appropriateness assessment (consultations with Korean institutions, surveys of historians, and focus groups with diverse community members), which will lead to an explicit go/no-go decision. In the event of a "no-go" decision, it will become a documented case study of respectful refusal. If approved, co-design workshops with Korean educators, language teachers, and museum professionals would develop culturally grounded use cases, followed by pilot implementation with Korean educational institutions, monthly community feedback sessions, and iterative refinement. The intellectual property model makes technical architecture open source while maintaining controlled access to the curated historical corpus, with any commercial applications supporting Korean historical research and cultural preservation. Critically, this framework makes refusal or modification as legitimate as approval (the goal is not to build a system at any cost), but only with genuine Korean partnership, where Korean communities exercise real authority over their heritage representation.
Recent research by Ted Underwood proved that AI models produce anachronisms even when trained exclusively on historical texts. Perfect historical authenticity is technically impossible. But rather than seeing this as a failure, we can make it educational.
When users learn to question AI outputs, "Why should I trust this? What evidence supports it? Whose interpretation is this?", they practice historical thinking. Our four-tier framework makes epistemic status visible, helping users distinguish between proven facts and scholarly interpretation. This is a core humanities skill that's often lost when information appears authoritative.
Systems should honestly state: "I am a computational model based on historical sources, not King Sejong himself." This acknowledgment models scholarly humility. Because the goal is not to tricka users into believing they are talking to the historical Sejong, it is helping them engage with Joseon Dynasty history in an accessible way while learning to think critically about sources and provide a scholarly example of transparency and intellectual integrity that can be adopted in other disciplines affected by the use of Generative AI.
Our understanding of this framework evolved through research and writing. We initially focused on technical sophistication but came to recognize that epistemic responsibility matters more than computational advancement. The system design itself is an interpretive act requiring ongoing reflection.
This models how digital humanities should approach AI development: iteratively and reflexively, acknowledging that frameworks embed cultural values that require continuous examination. There's no neutral implementation; every design choice reflects assumptions about what knowledge matters, how certainty should be represented, and whose authority counts.
The four-tier system works because it makes these interpretive choices explicit rather than hiding them. When the system classifies a response as "Tier 3: Reasonable Speculation," it's not just providing information; it's modeling historiographical reasoning. When it presents multiple contested views in Tier 4, it teaches users that historical interpretation involves debate, not just the accumulation of facts.
This approach resists what Drucker calls "digital positivism," the assumption that computational methods give us unmediated access to reality. Instead, it shows how digital tools can preserve the interpretive complexity that makes the humanities valuable while leveraging computational capabilities for accessibility and scale.
Creating AI systems that accurately represent historical figures responsibly requires the digital humanities to develop frameworks that preserve complexity rather than oversimplifying it. Our four-tier system shows how to distinguish computationally between documented facts, scholarly interpretations, reasonable guesses, and active debates. This operationalizes what historians already do, assessing certainty levels, tracking evidence, and acknowledging disagreement, in ways that AI systems can implement.
Success depends less on computational sophistication than on three key practices: maintaining a rigorous distinction between evidence and interpretation, honestly acknowledging uncertainty rather than hiding it, and granting cultural communities genuine authority over their heritage. For King Sejong, this means that Korean institutions will determine whether any AI representation is appropriate, a decision that external researchers cannot make, regardless of their technical capabilities. Our framework, therefore, incorporates Korean veto power from the outset, recognizing that the most responsible outcome may be community rejection of the entire project.
Our contribution to digital humanities demonstrates how computational tools can serve humanistic inquiry without oversimplifying it. The four-tier architecture offers a concrete approach to preserving interpretive complexity in AI systems. The challenges we identified, authenticity impossibility, and the need for cultural authority, represent ongoing concerns requiring continued attention, not problems with technical solutions.
This framework responds to what DH2026 calls "situated knowledge," referring to computational methods designed for specific cultural contexts, rather than universal templates. Our system works for King Sejong because it responds to Korean historiography, cultural values, and aligns with current scholarship. It might fail elsewhere, and that's appropriate. Good digital humanities design serves specific contexts rather than claiming universal applicability.
King Sejong created Hangul to democratize knowledge through linguistic innovation. A thoughtfully implemented AI system, developed through genuine Korean partnership and scholarly rigor, could extend that democratic vision, making Korean cultural heritage accessible while maintaining the historical accuracy and cultural respect that Sejong's memory deserves. But this can only happen through Korean leadership, not external prescription.