Score: 1

Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation

Published: April 8, 2025 | arXiv ID: 2504.05746v1

By: Zhihua Xu , Tianshui Chen , Zhijing Yang and more

Potential Business Impact:

Makes talking videos match speech perfectly.

Business Areas:

Motion Capture Media and Entertainment, Video

The paramount challenge in audio-driven One-shot Talking Head Animation (ADOS-THA) lies in capturing subtle imperceptible changes between adjacent video frames. Inherently, the temporal relationship of adjacent audio clips is highly correlated with that of the corresponding adjacent video frames, offering supplementary information that can be pivotal for guiding and supervising talking head animations. In this work, we propose to learn audio-visual correlations and integrate the correlations to help enhance feature representation and regularize final generation by a novel Temporal Audio-Visual Correlation Embedding (TAVCE) framework. Specifically, it first learns an audio-visual temporal correlation metric, ensuring the temporal audio relationships of adjacent clips are aligned with the temporal visual relationships of corresponding adjacent video frames. Since the temporal audio relationship contains aligned information about the visual frame, we first integrate it to guide learning more representative features via a simple yet effective channel attention mechanism. During training, we also use the alignment correlations as an additional objective to supervise generating visual frames. We conduct extensive experiments on several publicly available benchmarks (i.e., HDTF, LRW, VoxCeleb1, and VoxCeleb2) to demonstrate its superiority over existing leading algorithms.

Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos

Image and Video Processing

Makes talking videos smaller, clearer, and in sync.

16 Jun 2025 1

88%

Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings

Multimedia

Finds the perfect music for any video automatically.

6 Mar 2025 0

88%

Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation

CV and Pattern Recognition

Makes talking videos show real feelings, not fake ones.

25 Apr 2025 1

View PDF Login to Bookmark

Country of Origin

🇨🇳 China

Repos / Data Links

github.com

Page Count

11 pages

Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation

Makes talking videos match speech perfectly.

Technical Abstract

Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos

Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings

Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation