EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
By: Yixuan Zhang , Qing Chang , Yuxi Wang and more
Potential Business Impact:
Makes computer faces show real feelings when they talk.
Speech-driven 3D facial animation seeks to produce lifelike facial expressions that are synchronized with the speech content and its emotional nuances, finding applications in various multimedia fields. However, previous methods often overlook emotional facial expressions or fail to disentangle them effectively from the speech content. To address these challenges, we present EmoDiffusion, a novel approach that disentangles different emotions in speech to generate rich 3D emotional facial expressions. Specifically, our method employs two Variational Autoencoders (VAEs) to separately generate the upper face region and mouth region, thereby learning a more refined representation of the facial sequence. Unlike traditional methods that use diffusion models to connect facial expression sequences with audio inputs, we perform the diffusion process in the latent space. Furthermore, we introduce an Emotion Adapter to evaluate upper face movements accurately. Given the paucity of 3D emotional talking face data in the animation industry, we capture facial expressions under the guidance of animation experts using LiveLinkFace on an iPhone. This effort results in the creation of an innovative 3D blendshape emotional talking face dataset (3D-BEF) used to train our network. Extensive experiments and perceptual evaluations validate the effectiveness of our approach, confirming its superiority in generating realistic and emotionally rich facial animations.
Similar Papers
Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space
Multimedia
Makes cartoon faces show real feelings from sound.
RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis
CV and Pattern Recognition
Makes computer faces show real feelings from voices.
Audio Driven Real-Time Facial Animation for Social Telepresence
Graphics
Makes virtual people talk and move like real ones.