Daejeon, July 27–31
Figure 1: Juxtaposition of a single frame showing the original interview (left) and avatar recreation (right). Non-verbal features like eye-position and mouth movements are inferred from the original video.
Avatars offer communication that combines verbal and nonverbal expression with digital affordances such as control, reproducibility, and anonymity. While Deepfakes have showcased the risks of these technologies for deception and misinformation (Mirsky et al., 2021), avatars may also serve positive functions. They can provide a protected channel for self-disclosure on sensitive topics, allowing whistleblowers, witnesses, and people facing social stigma to share experiences without revealing their identity or sacrificing nonverbal expression. To use avatars as such a protective medium, we need to understand how audiences perceive them as communicative agents.
For DH and communications studies, this raises a central question of engagement with online media: how do viewers respond to emotional messages delivered by artificial agents instead of human speakers?
We addressed this question with an online experiment in Germany. Participants were randomly assigned to watch a short interview including a person’s emotional self-disclosure on a stigmatized topic (sex work). The control group viewed the original human interviewee, while the experimental group viewed an AI-generated avatar reproducing her verbal and nonverbal behaviour (see Figure 1). We compared message acceptance, prosocial attitudes towards the stigmatized group, perceived authenticity, and empathy. The latter two were considered mediating variables that might explain effects on message acceptance and attitudes.
We present preliminary results and a Video-to-Video inference pipeline that enabled us to generate these tightly controlled stimuli.
Research on audiences’ empathic and prosocial responses to avatars is limited. A study by Roth et al. (2019) on self-disclosure through avatars found that such disclosures can be perceived as authentic and elicit similar levels of empathy and prosocial intentions compared to a human actor. However, their human stimulus was a lay actor, not an authentic testimonial of a person affected by the issue. Thus, it remains unclear how the avatar would have been perceived in comparison with an authentic testimonial. Moreover, their pipeline required complex and costly hardware and could not process pre-existing audiovisual material.
The current ecosystem of generative video tools is fragmented. Commercial avatar platforms produce convincing synthetic presenters, but are costly and enforce strict content policies, making them unusable for stigmatized topics. Open-source one-in-all avatar solutions often lack stable faces, liveliness, or convincing lip synchronization.
Digital human research has advanced in facial synthesis (Menze et al., 2025) and motion generation (Chen et al., 2024), but typically requires complex simulation and expensive hardware. Still, they introduce subtle artifacts or exaggerated expressions. For perception experiments, such issues risk becoming confounding variables. This aligns with DH concerns about how artificial representations shape interpretation, trust, and engagement.
Our stimuli are based on a publicly available interview format in which individuals with stigmatized experiences explain how they are affected and answer audience questions. We selected an interview with a former forced prostitute.
All participants received the same verbal content. Only facial and vocal speaker identity differed (face conversion with preserved nonverbal expression, voice conversion with preserved prosody).
Our central requirement for the avatar condition was to create a second video that:
This exceeds the ability of traditional reenactments with actors, as they cannot satisfy constraints 1 & 3.
Figure 2: Iterations of creating the avatar (left) that closely matches the original appearance (right).
To meet these requirements, we developed a video-driven avatar-generation pipeline adopting different generative AI steps:
The resulting avatar mirrors the original interviewee at the level of communication behaviour. The same pipeline allows systematic experimental variation of facial and vocal features without re-recording and while keeping patterns of nonverbal expression constant. This demonstrates how open-source tools can form an infrastructure for experiments on engagement with AI-generated media.
Preliminary Structural-Equation-Modeling (SEM) (Kline, 2011) suggests that, compared to the original interview, using an avatar significantly reduced perceived authenticity and, thus, empathic engagement. However, levels of message acceptance and prosocial attitudes did not significantly differ. Thus, the net effect on message acceptance and prosocial attitudes was neutral, suggesting that avatars can be a viable mode for self-disclosure on stigmatized topics while protecting privacy. Nevertheless, we observed a negative indirect effect of the avatar on message acceptance and prosocial attitudes via reduced authenticity and empathy, indicating that avatars should be used with caution and further refinement is needed.
This paper reported preliminary results from an online-experiment comparing audience responses to humans and avatars. We matched an existing, real-world interview with a virtual avatar mirroring the original interviewee’s verbal and non-verbal communication patterns and presented an open-source avatar-generation pipeline that reduces costs and hardware requirements.
Video-to-Video inference allowed us to manipulate identifying facial features while keeping patterns of nonverbal communication constant. This method also allowed us to use pre-existing content instead of hardware-intensive simulation. The self-hostable open-source tools also support work with sensitive topics typically blocked by third-party services.
Beyond our pipeline and preliminary results, we also emphasize potentials for future engagement studies and long-term investigations on how audiences respond to AI-generated avatars as AI content becomes more common.
The authors acknowledge the financial support by the Federal Ministry of Education and Research of Germany and by Sächsische Staatsministerium für Wissenschaft, Kultur und Tourismus in the programme Center of Excellence for AI-research „Center for Scalable Data Analytics and Artificial Intelligence Dresden/Leipzig“, project identification number: ScaDS.AI.