Score: 0

MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation

Published: December 20, 2025 | arXiv ID: 2512.18181v1

By: Kaixing Yang , Jiashu Zhu , Xulong Tang and more

With the rise of online dance-video platforms and rapid advances in AI-generated content (AIGC), music-driven dance generation has emerged as a compelling research direction. Despite substantial progress in related domains such as music-driven 3D dance generation, pose-driven image animation, and audio-driven talking-head synthesis, existing methods cannot be directly adapted to this task. Moreover, the limited studies in this area still struggle to jointly achieve high-quality visual appearance and realistic human motion. Accordingly, we present MACE-Dance, a music-driven dance video generation framework with cascaded Mixture-of-Experts (MoE). The Motion Expert performs music-to-3D motion generation while enforcing kinematic plausibility and artistic expressiveness, whereas the Appearance Expert carries out motion- and reference-conditioned video synthesis, preserving visual identity with spatiotemporal coherence. Specifically, the Motion Expert adopts a diffusion model with a BiMamba-Transformer hybrid architecture and a Guidance-Free Training (GFT) strategy, achieving state-of-the-art (SOTA) performance in 3D dance generation. The Appearance Expert employs a decoupled kinematic-aesthetic fine-tuning strategy, achieving state-of-the-art (SOTA) performance in pose-driven image animation. To better benchmark this task, we curate a large-scale and diverse dataset and design a motion-appearance evaluation protocol. Based on this protocol, MACE-Dance also achieves state-of-the-art performance. Project page: https://macedance.github.io/

DanceMosaic: High-Fidelity Dance Generation with Multimodal Editability

Graphics

Creates realistic, editable 3D dances from music and text.

6 Apr 2025 1

89%

DanceMeld: Unraveling Dance Phrases with Hierarchical Latent Codes for Music-to-Dance Synthesis

Other Computer Science

Makes computers create realistic dance moves from music.

30 Nov 2023 0

89%

Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation

CV and Pattern Recognition

Makes dancing robots move to music perfectly.

12 Dec 2025 0

View PDF Login to Bookmark

MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation

Technical Abstract

DanceMosaic: High-Fidelity Dance Generation with Multimodal Editability

DanceMeld: Unraveling Dance Phrases with Hierarchical Latent Codes for Music-to-Dance Synthesis

Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation