Models & pricingModelsGLM-5.2-Vision-NVFP4
GL

GLM-5.2-Vision-NVFP4

baseten

open weights·Active·multimodal general·Updated Jul 20, 2026

GLM-5.2-Vision-NVFP4 is a 381.0B-parameter image text to text model developed by baseten. Built on the Glm5vForConditionalGeneration architecture using sglang. Released on 2026-07-20 with 126 likes and 2,276 downloads on Hugging Face.

GLM-5.2-Vision-NVFP4

Model Overview

GLM-5.2-Vision-NVFP4 is a 381.0B-parameter image text to text model developed by baseten. Built on the Glm5vForConditionalGeneration architecture. Released on 2026-07-20.


📊 Quick Specs

Specification Table
SpecificationValue
Parameters381.0B
ArchitectureGlm5vForConditionalGeneration
Taskimage text to text
Modalitytext, image
LicenseMIT
Frameworksglang
MoENo
Languages

✨ Key Features

  • 381.0B parameters
  • Built on Glm5vForConditionalGeneration architecture (sglang)
  • Primary task: image text to text (text, image modality)
  • Open-weights under MIT license — self-hostable and fine-tunable

📈 Community Adoption

  • 126 likes on Hugging Face
  • 2,276 downloads on Hugging Face

🔗 Resources


📜 License & Access

MIT — Open-weights model available for download, fine-tuning, and self-hosted deployment.

Related models comparison