Daejeon, July 27–31
Poetry translation is inherently a negotiation with loss. While Jakobson (2021) conceptualizes translation as the conversion of linguistic signs, the affective resonance and cultural density embedded within classical poetry frequently dissipate across disparate linguistic boundaries. Furthermore, although scholars like Forster (2023) attribute unique pictorial qualities to Chinese poetry based on its ideographic roots, we reject this premise as a flawed framework. Historical evolution clearly demonstrates that Chinese characters have transitioned into profound abstraction. Accordingly, the "image gap" addressed in this study does not emerge from the visual morphology of the characters, but rather from untranslatable cultural and affective metaphors. This complication is especially pronounced in the Shijing (Book of Songs), where the aesthetic tension of Ancient Chinese—driven by parataxis and intentional silence—is routinely flattened by the hypotactic constraints of English syntax (Yan 2020).
To theoretically ground our research and situate it within contemporary advancements, this study incorporates emotion mining. Although this field has successfully decoded emotional phenomena in modern linguistic and visual domains (e.g., social media sentiment analysis and affective image generation), its integration into ancient, cross-modal poetic contexts remains largely unexplored. To address this deficiency, we propose a "Cross-Modal Affective Compensation Mechanism." We hypothesize that when linguistic translation reaches the absolute limits of "untranslatability" (Ping 1999), multimodal Large Language Models (LLMs) can effectively bridge the divide by translating implicit textual affect into explicit visual atmospheres.
Our approach shifts the paradigm from semantic equivalence to affective compensation. We recognize that translating a poem is not merely swapping words, but rather reconstructing the "text-emotion-image" triad. Thus, our pipeline is designed not to produce literal illustrations, but to computationally visualize unstated emotions using image generation techniques.
We view the Book of Songs as a high-compression cultural archive requiring affective decompression. As illustrated in Figure 1, our method follows three stages:
Figure 1. Three-Stage Experimental Framework.
STAGE 1: Narrative Expansion & Contextual Reconstruction
We use Xunzi-Qwen3-8B, a model fine-tuned on ancient Chinese texts (XunziALLM Team 2025). Acting as a cultural decoder, it expands concise ancient verses into rich English narratives. This step prioritizes emotional context over literal accuracy. To achieve this, we utilize specific prompt templates, as shown in Figure 2.
Figure 2. Illustration of the Affective Decoding Prompt.
STAGE 2: Multi-modal Affective Infusion
Using Nano Banana Pro, an image generation model based on the Gemini 3 architecture (Google 2025a), we generate visual representations from the text employing a responsive information-density strategy (Wu et al. 2025). We compare four prompting conditions:
(1)Group A: Ancient Original
(2)Group B: Literal Translation
(3)Group C: Narrative Expansion
(4)Group D: Zero-Emotion Scientific Baseline
As demonstrated in Figure 3, this comparative grouping visualizes the "loss" of sentiment in literal translation versus the "gain" achieved through narrative expansion.
Figure 3. Example of Text Affective Visualization Generation.
STAGE 3: Bidirectional Cross-Modal Affective Evaluation
Given that poetic emotion is profoundly subjective, our study does not seek to establish absolute objective quantification, but rather a persuasive framework of relative alignment. We employ Gemini 3.1 Pro for bidirectional evaluation (Google 2025b). First, we conduct a "visual Turing test" to rate visual quality and emotional clarity (Zhao et al. 2018; Hessel et al. 2021). Second, we prompt Gemini 3.1 Pro to perform Affective Backtracking, extracting emotional keywords from the generated images to test whether the visual output successfully conveys the poem's core emotion without exposure to the original text.
Our analysis culminates in the proposal of the Semantic-Affective Trade-off Model. As illustrated in the scatter plot (Figure 4), the x-axis maps Semantic Precision (measured by object-matching metrics such as CLIPScore, scaled 0 to 1), while the y-axis delineates Affective Fidelity (based on normalized, LLM-evaluated Likert ratings of emotional resonance).
Figure 4. The Semantic-Affective Trade-off Model.
To anchor this model in textual practice, we examine the Shijing verse Jian Jia (Reeds and Rushes). The literal translation baseline (Group B) yields high semantic precision—faithfully rendering the riparian flora—yet suffers from an affective deficit. The generated scenes appear sterile and emotionally vacant because the underlying motif of unfulfilled longing eludes literal translation. Conversely, our narrative expansion protocol (Group C) forces the LLM to manifest this subtext through atmospheric cues such as "receding mist." This intervention successfully translates abstract longing into tangible visual mood, situating Group C firmly within the "Poetic Zone" (High Affect/Moderate Semantics).
Beyond methodological insights, this research challenges the tendency of AI to default to Western-centric cultural norms (Tao et al. 2024; Liu 2025). Instead, we observe the "Book of Songs Effect," where Multimodal LLMs help universalize specific cultural imagery. By converting culturally specific symbols into a universal visual vocabulary (Oppenlaender 2022), the AI enables global audiences to access the emotional core of the ancient text. This aligns with recent findings that generative AI can significantly enhance visitor engagement and cultural dissemination (Fu et al. 2024).
We demonstrate that Generative AI is not just a translation tool, but an affective medium capable of crossing language boundaries. Furthermore, the project presentation will incorporate QR codes to facilitate real-time access to the 'Semantic-Affective' dashboard. Attendees can view side-by-side image generations (Literal vs. Narrative) for all poem segments. We will also demonstrate the "Backward Validation" process live, inviting attendees to guess the emotion of a generated image before revealing the original Shijing verse, replicating our methodology in real-time.