SCAIL 2: End-to-End In-Context Controlled Character Animation
Model Overview
SCAIL 2 is an open-source, end-to-end framework for controlled character animation and video-to-video motion transfer developed by researchers at Tsinghua University (KEG Group) and Z.ai (Wenhao Yan, Fengjia Guo, Jie Tang, et al.).
It enables video-to-video motion transfer directly from a driving video to a reference character image/video without relying on intermediate representations such as 3D skeletons, OpenPose maps, depth maps, or foreground masks.
Key Features
- End-to-End Motion Transfer Architecture: Bypasses traditional pose skeleton/depth map extraction, eliminating information loss and allowing motion transfer onto arbitrary character styles and non-humanoids.
- Unified In-Context Mask & Mode-Specific RoPE: Uses in-context mask conditioning alongside Rotary Position Embeddings (RoPE) to unify single-character animation, character replacement, and multi-character interaction.
- MotionPair-60K Dataset: Trained on a heterogeneous dataset of 60,000 synthetic motion transfer pairs.
- Bias-Aware Direct Preference Optimization (DPO): Post-training DPO improves generation quality in fine-detail regions like fingers, facial expressions, and limb overlaps.
- ComfyUI Integration: Features native integration and memory-efficient offloading for ComfyUI.
Verified Project Links
- Project Website: https://teal024.github.io/SCAIL-2/
- arXiv Paper: https://arxiv.org/abs/2606.10804
- GitHub Repository: https://github.com/zai-org/SCAIL-2
- Hugging Face Model: https://huggingface.co/zai-org/SCAIL-2
Benchmarks & Evaluation
- Evaluated on TikTok and X-Dance datasets for pose & motion consistency.
- Studio-Bench: Achieved >70% win-rate over skeleton-based baselines in real-world cross-identity animation.
Key Features
End-to-End Motion Transfer Architecture: Bypasses traditional pose skeleton/depth map extraction, enabling motion transfer onto arbitrary character styles
Unified In-Context Mask & Mode-Specific RoPE Conditioning: Uses in-context mask conditioning and spatial/temporal RoPE to unify single-character and multi-character interaction
MotionPair-60K Dataset: Trained on 60,000 synthetic motion transfer pairs constructed via reverse-driving pipeline
Bias-Aware Direct Preference Optimization (DPO): Mitigates synthetic distribution bias, improving fine details in fingers and facial expressions
Emergent Zero-Shot Generalization & ComfyUI Support: Strong zero-shot generalization to non-humanoid/animal animation and native ComfyUI integration
You might also want to compare
Verified Sources
Tags
Model Specs
Parameters
Undisclosed
Context Window
undisclosed
License
Apache-2.0
Deployment
Resources & Links
Curator Notes
Verified paper arXiv:2606.10804 and open-source project from Tsinghua University KEG Group and Z.ai.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model