Back to Wan
Open WeightsVideo GenaudiovideoUpdated July 14, 2026

Wan-Dancer: Hierarchical Framework for Minute-scale Music-to-Dance Generation

Model Overview

Wan-Dancer is a 14-billion parameter open-weights hierarchical framework developed by Tongyi Lab at Alibaba Group for generating long-duration, high-quality, rhythmically synchronized dance videos from music audio and a reference character image. Wan-Dancer achieves minute-scale coherent dance generation — far exceeding the short-clip limitations of prior music-to-dance systems — while supporting multiple diverse dance genres and customizable reference-based personalization.


🎶 Dance Genres Supported

Specification Table
GenreDescription
Chinese ClassicalTraditional fluid body movements with cultural stylization
K-popEnergetic idol-style choreography synchronized to pop music
Street DanceBreaking, popping, locking, and freestyle urban dance
TapPrecise foot percussion movements tied to rhythmic beats
LatinSalsa, bachata, and rumba-style partner dance patterns

✨ Key Features

Specification Table
FeatureDescription
Minute-scale Coherent GenerationGenerates long-duration dance videos with maintained global choreographic structure beyond short-clip baselines
Hierarchical FrameworkSeparates global choreography planning from local motion rendering for temporally consistent long-form generation
Music-Driven SynchronizationAnalyzes musical beat, rhythm, and mood to synchronize body movements with audio
Customizable ReferenceAccepts arbitrary character reference images; supports keyframe control for outfit changes and movement customization
Multi-Reference CompositionSupports single music + multiple reference characters or single character + multiple music tracks

🔗 Resources


👥 Authors & Institution

Developed by the Tongyi Lab at Alibaba Group:

  • Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Ruoshi Zhang, Yi Lu, Gang Cheng, Bang Zhang

Institution: Tongyi Lab, Alibaba Group Model Size: 14B parameters (open-weights)

Key Features

Minute-scale Coherent Generation: Generates long-duration dance videos (minute-scale) with maintained global structure and temporal continuity beyond existing short-clip limitations

Feature 01

Cross-Genre Dance Synthesis: Supports diverse dance styles including Chinese Classical, K-pop, Street, Tap, and Latin with genre-specific rhythmic characteristics

Feature 02

Hierarchical Framework: Uses a two-stage hierarchical approach separating global choreography planning from local motion rendering for coherent long-form generation

Feature 03

Music-Driven Synchronization: Analyzes musical beat, rhythm, and mood to synchronize body movements and expressions with audio in real-time

Feature 04

Customizable Reference: Accepts arbitrary character reference images and supports keyframe control for outfit changes and movement customization

Feature 05

You might also want to compare

Verified Sources

Tags

music-to-dancevideo-generationdance-synthesisaudio-drivenlong-form-video

Model Specs

open-weights

Parameters

14B

Context Window

unknown

License

Other/Custom

Deployment

self-hostableapi-only

Resources & Links

Lineage

Model Family

Part of the Wan family

Only release in this line currently tracked.

Curator Notes

14B parameter open-weights model by Tongyi Lab (Alibaba). Authors: Mingyang Huang, Peng Zhang, Li Hu et al. Supports Chinese Classical, K-pop, Street, Tap, Latin dance genres.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model