GL
GLM-5.2-Vision-NVFP4
baseten
ID: glm-52-vision-nvfp4
open weights·Active·multimodal general·Updated Jul 20, 2026
GLM-5.2-Vision-NVFP4 is a 381.0B-parameter image text to text model developed by baseten. Built on the Glm5vForConditionalGeneration architecture using sglang. Released on 2026-07-20 with 126 likes and 2,276 downloads on Hugging Face.
GLM-5.2-Vision-NVFP4
Model Overview
GLM-5.2-Vision-NVFP4 is a 381.0B-parameter image text to text model developed by baseten. Built on the Glm5vForConditionalGeneration architecture. Released on 2026-07-20.
📊 Quick Specs
Specification Table
| Specification | Value |
|---|---|
| Parameters | 381.0B |
| Architecture | Glm5vForConditionalGeneration |
| Task | image text to text |
| Modality | text, image |
| License | MIT |
| Framework | sglang |
| MoE | No |
| Languages | — |
✨ Key Features
- 381.0B parameters
- Built on Glm5vForConditionalGeneration architecture (sglang)
- Primary task: image text to text (text, image modality)
- Open-weights under MIT license — self-hostable and fine-tunable
📈 Community Adoption
- 126 likes on Hugging Face
- 2,276 downloads on Hugging Face
🔗 Resources
- Hugging Face Hub: GLM-5.2-Vision-NVFP4 on Hugging Face
📜 License & Access
MIT — Open-weights model available for download, fine-tuning, and self-hosted deployment.