VidiHand
Model Overview
VidiHand is a undisclosed-parameter model developed by Academic/Research. Released on 2026-06-29.
📊 Quick Specs
| Specification | Value |
|---|---|
| Parameters | undisclosed |
| Task | other |
| Modality | video, 3d |
| License | Other/Custom |
| Type | open-weights |
✨ Key Features
- Detector-Free Full-Frame Processing without localized cropping or hand detectors
- Leverages Internet-Scale Pretrained Video Diffusion Models (Wan2.1-VACE)
- Extreme Robustness to Heavy Hand-Object and Hand-Hand Occlusions
- Eliminates Test-Time Optimization (TTO) and post-hoc temporal infilling
- State-of-the-Art Temporal Smoothness (jitter down to 3.18 mm/frame) and Pose Accuracy (21.668 mm MPJPE-p on ARCTIC)
🔗 Resources
- Hugging Face: VidiHand
- GitHub: Repository
- Paper: arXiv
- Website: Project Page
📜 License & Access
Other/Custom — See repository for specific license details.
Key Features
Detector-Free Full-Frame Processing without localized cropping or hand detectors
Leverages Internet-Scale Pretrained Video Diffusion Models (Wan2.1-VACE)
Extreme Robustness to Heavy Hand-Object and Hand-Hand Occlusions
Eliminates Test-Time Optimization (TTO) and post-hoc temporal infilling
State-of-the-Art Temporal Smoothness (jitter down to 3.18 mm/frame) and Pose Accuracy (21.668 mm MPJPE-p on ARCTIC)
You might also want to compare
Verified Sources
Tags
Model Specs
Parameters
Undisclosed
Context Window
undisclosed
License
Other/Custom
Deployment
Resources & Links
Curator Notes
Partially enriched via migration on 2026-07-25. Manual review recommended.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model