MA
Mage-VL
microsoft
ID: mage-vl
open weights·Active·multimodal general·Updated Jul 25, 2026
Mage-VL is a 4.7B-parameter image text to text model developed by microsoft. Built on the MageVLForConditionalGeneration architecture using transformers. Released on 2026-07-25 with 121 likes and 2,951 downloads on Hugging Face.
Mage-VL
Model Overview
Mage-VL is a 4.7B-parameter image text to text model developed by microsoft. Built on the MageVLForConditionalGeneration architecture. Released on 2026-07-25.
📊 Quick Specs
Specification Table
| Specification | Value |
|---|---|
| Parameters | 4.7B |
| Architecture | MageVLForConditionalGeneration |
| Task | image text to text |
| Modality | text, image |
| License | APACHE-2.0 |
| Framework | transformers |
| MoE | No |
| Languages | — |
✨ Key Features
- 4.7B parameters
- Built on MageVLForConditionalGeneration architecture (transformers)
- Primary task: image text to text (text, image modality)
- Open-weights under APACHE-2.0 license — self-hostable and fine-tunable
📈 Community Adoption
- 121 likes on Hugging Face
- 2,951 downloads on Hugging Face
🔗 Resources
- Hugging Face Hub: Mage-VL on Hugging Face
- Paper: arXiv
📜 License & Access
APACHE-2.0 — Open-weights model available for download, fine-tuning, and self-hosted deployment.