Wan-Dancer
Wan-Dancer: Hierarchical Framework for Minute-scale Music-to-Dance Generation
Model Overview
Wan-Dancer is a 14-billion parameter open-weights hierarchical framework developed by Tongyi Lab at Alibaba Group for generating long-duration, high-quality, rhythmically synchronized dance videos from music audio and a reference character image. Wan-Dancer achieves minute-scale coherent dance generation — far exceeding the short-clip limitations of prior music-to-dance systems — while supporting multiple diverse dance genres and customizable reference-based personalization.
🎶 Dance Genres Supported
| Genre | Description |
|---|---|
| Chinese Classical | Traditional fluid body movements with cultural stylization |
| K-pop | Energetic idol-style choreography synchronized to pop music |
| Street Dance | Breaking, popping, locking, and freestyle urban dance |
| Tap | Precise foot percussion movements tied to rhythmic beats |
| Latin | Salsa, bachata, and rumba-style partner dance patterns |
✨ Key Features
| Feature | Description |
|---|---|
| Minute-scale Coherent Generation | Generates long-duration dance videos with maintained global choreographic structure beyond short-clip baselines |
| Hierarchical Framework | Separates global choreography planning from local motion rendering for temporally consistent long-form generation |
| Music-Driven Synchronization | Analyzes musical beat, rhythm, and mood to synchronize body movements with audio |
| Customizable Reference | Accepts arbitrary character reference images; supports keyframe control for outfit changes and movement customization |
| Multi-Reference Composition | Supports single music + multiple reference characters or single character + multiple music tracks |
🔗 Resources
| Resource | Link |
|---|---|
| Project Page | humanaigc.github.io/wan-dancer-project |
| arXiv Paper | arxiv.org/abs/2607.09581 |
| GitHub | github.com/Wan-Video/Wan-Dancer |
| HuggingFace Model | huggingface.co/Wan-AI/Wan-Dancer-14B |
| ModelScope Demo | modelscope.ai/studios/Wan-AI/Wan-Dancer |
👥 Authors & Institution
Developed by the Tongyi Lab at Alibaba Group:
- Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Ruoshi Zhang, Yi Lu, Gang Cheng, Bang Zhang
Institution: Tongyi Lab, Alibaba Group Model Size: 14B parameters (open-weights)
Key Features
Minute-scale Coherent Generation: Generates long-duration dance videos (minute-scale) with maintained global structure and temporal continuity beyond existing short-clip limitations
Cross-Genre Dance Synthesis: Supports diverse dance styles including Chinese Classical, K-pop, Street, Tap, and Latin with genre-specific rhythmic characteristics
Hierarchical Framework: Uses a two-stage hierarchical approach separating global choreography planning from local motion rendering for coherent long-form generation
Music-Driven Synchronization: Analyzes musical beat, rhythm, and mood to synchronize body movements and expressions with audio in real-time
Customizable Reference: Accepts arbitrary character reference images and supports keyframe control for outfit changes and movement customization
You might also want to compare
Verified Sources
Tags
Model Specs
Parameters
14B
Context Window
unknown
License
Other/Custom
Deployment
Resources & Links
Lineage
Model Family
Part of the Wan family
Only release in this line currently tracked.
Curator Notes
14B parameter open-weights model by Tongyi Lab (Alibaba). Authors: Mingyang Huang, Peng Zhang, Li Hu et al. Supports Chinese Classical, K-pop, Street, Tap, Latin dance genres.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model