Score: 3

AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control

Published: November 26, 2025 | arXiv ID: 2511.21146v1

By: Xinyue Guo , Xiaoran Yang , Lipan Zhang and more

BigTech Affiliations: Xiaomi

Potential Business Impact:

Changes video sounds using pictures and words.

Business Areas:

Video Editing Content and Publishing, Media and Entertainment, Video

Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and suboptimal audio quality. To address this, we propose AV-Edit, a generative sound effect editing framework that enables fine-grained editing of existing audio tracks in videos by jointly leveraging visual, audio, and text semantics. Specifically, the proposed method employs a specially designed contrastive audio-visual masking autoencoder (CAV-MAE-Edit) for multimodal pre-training, learning aligned cross-modal representations. These representations are then used to train an editorial Multimodal Diffusion Transformer (MM-DiT) capable of removing visually irrelevant sounds and generating missing audio elements consistent with video content through a correlation-based feature gating training strategy. Furthermore, we construct a dedicated video-based sound editing dataset as an evaluation benchmark. Experiments demonstrate that the proposed AV-Edit generates high-quality audio with precise modifications based on visual content, achieving state-of-the-art performance in the field of sound effect editing and exhibiting strong competitiveness in the domain of audio generation.

Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

Multimedia

Makes videos and sounds match perfectly.

8 Dec 2025 1

91%

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model

Sound

Changes sounds in audio using text instructions.

23 Dec 2025 0

91%

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model

Sound

Edits sounds in audio using text instructions.

23 Dec 2025 0

View PDF Login to Bookmark

Country of Origin

🇨🇳 China

Repos / Data Links

github.com

Page Count

14 pages

AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control

Changes video sounds using pictures and words.

Technical Abstract

Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model