DH 2026

Daejeon, July 27–31

Fri, July 3109:00–10:30S049206-208
Short Paper

Experimental Evidence on Perceptions of Avatars in Stigmatized Communication Contexts

Richard Khulusi
Image and Signal Processing Group, Leipzig University; Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI) Dresden/Leipzig, Leipzig University, Germany · khulusi@uni-leipzig.de
Simon Luebke
Institute of Communication and Media Studies, Leipzig University · simon.luebke@uni-leipzig.de
Christal Bürgel
Institute of Communication and Media Studies, Leipzig University · christal.buergel@uni-leipzig.de
Gerik Scheuermann
Image and Signal Processing Group, Leipzig University; Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI) Dresden/Leipzig, Leipzig University, Germany · scheuermann@informatik.uni-leipzig.de
Anne Bartsch
Institute of Communication and Media Studies, Leipzig University · anne.bartsch@uni-leipzig.de

Figure 1: Juxtaposition of a single frame showing the original interview (left) and avatar recreation (right). Non-verbal features like eye-position and mouth movements are inferred from the original video.

Introduction

Avatars offer communication that combines verbal and nonverbal expression with digital affordances such as control, reproducibility, and anonymity. While Deepfakes have showcased the risks of these technologies for deception and misinformation (Mirsky et al., 2021), avatars may also serve positive functions. They can provide a protected channel for self-disclosure on sensitive topics, allowing whistleblowers, witnesses, and people facing social stigma to share experiences without revealing their identity or sacrificing nonverbal expression. To use avatars as such a protective medium, we need to understand how audiences perceive them as communicative agents.

For DH and communications studies, this raises a central question of engagement with online media: how do viewers respond to emotional messages delivered by artificial agents instead of human speakers?

We addressed this question with an online experiment in Germany. Participants were randomly assigned to watch a short interview including a person’s emotional self-disclosure on a stigmatized topic (sex work). The control group viewed the original human interviewee, while the experimental group viewed an AI-generated avatar reproducing her verbal and nonverbal behaviour (see Figure 1). We compared message acceptance, prosocial attitudes towards the stigmatized group, perceived authenticity, and empathy. The latter two were considered mediating variables that might explain effects on message acceptance and attitudes.

We present preliminary results and a Video-to-Video inference pipeline that enabled us to generate these tightly controlled stimuli.

Communication Perspective

Research on audiences’ empathic and prosocial responses to avatars is limited. A study by Roth et al. (2019) on self-disclosure through avatars found that such disclosures can be perceived as authentic and elicit similar levels of empathy and prosocial intentions compared to a human actor. However, their human stimulus was a lay actor, not an authentic testimonial of a person affected by the issue. Thus, it remains unclear how the avatar would have been perceived in comparison with an authentic testimonial. Moreover, their pipeline required complex and costly hardware and could not process pre-existing audiovisual material.

Technical and DH Perspectives

The current ecosystem of generative video tools is fragmented. Commercial avatar platforms produce convincing synthetic presenters, but are costly and enforce strict content policies, making them unusable for stigmatized topics. Open-source one-in-all avatar solutions often lack stable faces, liveliness, or convincing lip synchronization.

Digital human research has advanced in facial synthesis (Menze et al., 2025) and motion generation (Chen et al., 2024), but typically requires complex simulation and expensive hardware. Still, they introduce subtle artifacts or exaggerated expressions. For perception experiments, such issues risk becoming confounding variables. This aligns with DH concerns about how artificial representations shape interpretation, trust, and engagement.

Data and Stimulus Construction

Our stimuli are based on a publicly available interview format in which individuals with stigmatized experiences explain how they are affected and answer audience questions. We selected an interview with a former forced prostitute.

All participants received the same verbal content. Only facial and vocal speaker identity differed (face conversion with preserved nonverbal expression, voice conversion with preserved prosody).

Our central requirement for the avatar condition was to create a second video that:

  • Preserves key nonverbal features
  • Protects interviewees’ identity
  • Is reproducible and adaptable

This exceeds the ability of traditional reenactments with actors, as they cannot satisfy constraints 1 & 3.

Figure 2: Iterations of creating the avatar (left) that closely matches the original appearance (right).

Avatar Generation Pipeline

To meet these requirements, we developed a video-driven avatar-generation pipeline adopting different generative AI steps:

  • Portrait Creation: Creation of a portrait resembling the interviewee, but without identifying facial features (Figure 2) using Fooocus (R2, R3).
  • Idle-Motion Generation: Transformation of the static image to a short idle-movement clip using SkyReel (Chen et al., 2025) adding minimal movement and posture shift to create liveliness.
  • Video-to-Video Inference: Inferring detailed movement action from the original video to the idle-animation, creating a close movement double with LivePortrait (Guo et al., 2024), preserving head motions, facial expressions, and lip movement with the original’s timing and expressiveness (see Figure 1) (R1, R3).
  • Voice Modification: Voice-to-Voice conversion using ElevenLabs, adapting the voice to anonymise the interviewee’s vocal features (R2), while keeping linguistic content and prosodic contour (R1).

The resulting avatar mirrors the original interviewee at the level of communication behaviour. The same pipeline allows systematic experimental variation of facial and vocal features without re-recording and while keeping patterns of nonverbal expression constant. This demonstrates how open-source tools can form an infrastructure for experiments on engagement with AI-generated media.

Preliminary Results

Preliminary Structural-Equation-Modeling (SEM) (Kline, 2011) suggests that, compared to the original interview, using an avatar significantly reduced perceived authenticity and, thus, empathic engagement. However, levels of message acceptance and prosocial attitudes did not significantly differ. Thus, the net effect on message acceptance and prosocial attitudes was neutral, suggesting that avatars can be a viable mode for self-disclosure on stigmatized topics while protecting privacy. Nevertheless, we observed a negative indirect effect of the avatar on message acceptance and prosocial attitudes via reduced authenticity and empathy, indicating that avatars should be used with caution and further refinement is needed.

Conclusion

This paper reported preliminary results from an online-experiment comparing audience responses to humans and avatars. We matched an existing, real-world interview with a virtual avatar mirroring the original interviewee’s verbal and non-verbal communication patterns and presented an open-source avatar-generation pipeline that reduces costs and hardware requirements.

Video-to-Video inference allowed us to manipulate identifying facial features while keeping patterns of nonverbal communication constant. This method also allowed us to use pre-existing content instead of hardware-intensive simulation. The self-hostable open-source tools also support work with sensitive topics typically blocked by third-party services.

Beyond our pipeline and preliminary results, we also emphasize potentials for future engagement studies and long-term investigations on how audiences respond to AI-generated avatars as AI content becomes more common.

Acknowledgements

The authors acknowledge the financial support by the Federal Ministry of Education and Research of Germany and by Sächsische Staatsministerium für Wissenschaft, Kultur und Tourismus in the programme Center of Excellence for AI-research „Center for Scalable Data Analytics and Artificial Intelligence Dresden/Leipzig“, project identification number: ScaDS.AI.

References
  1. Chen, Z., Cao, J., Chen, Z., Li, Y., Ma, C. (2024). EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditioning. arXiv preprint arXiv:2407.08136
  2. Chen, G., Lin, D., Yang, J., Lin, C., Zhu, J.,  Fan, M., Zhang, H., Chen, S., Chen, Z., Ma, C., and others (2025). Skyreels-v2: Infinite-length film generative model. arXiv preprint arXiv:2504.13074.
  3. Guo, J., Zhang, D., Liu, X., Zhong, Z., Zhang, Y., Wan, P., & Zhang, D. (2024). Liveportrait: Efficient portrait animation with stitching and retargeting control. arXiv preprint arXiv:2407.03168.
  4. Kline, R. B. (2011). Principles and practice of structural equation modeling (3rd ed.). The Guilford Press.
  5. Menzel T, Wolf E, Wenninger S, Spinczyk N, Holderrieth L, Wienrich C, Schwanecke U, Latoschik ME and Botsch M. (2025). Avatars for the masses: smartphone-based reconstruction of humans for virtual reality. Front. Virtual Real.
  6. Mirsky, Y., & Lee, W. (2021). The creation and detection of deepfakes: A survey. ACM computing surveys 
  7. Roth, D., Bloch, C., Schmitt, J., Frischlich, L., Latoschik, M.E., Bente, G. (2019). Perceived Authenticity, Empathy, and Pro-social Intentions evoked through Avatar-mediated Self-disclosures. Proceedings of Mensch Und Computer 2019