Models & pricing / Models / DeepSeek-V4-Flash
DeepSeek-V4-Flash Overview
DeepSeek-V4-Flash is a 284B-parameter Mixture-of-Experts (MoE) text generation model developed by deepseek-ai. Built on the DeepseekV4ForCausalLM architecture using transformers, it activates 13B parameters per token. Released on 2026-04-22 with 1,945 likes and 2,814,414 downloads on Hugging Face (where compressed file configurations cause systems to display it as 158.1B).
Choosing a model
If you're unsure which model to use, start with DeepSeek-V4-Flash for complex agentic coding, reasoning, and enterprise workloads. For lightweight, low-latency autocomplete or edge tasks, consider smaller parameters.
All current deepseek-ai models support text and multimodal input, multilingual reasoning, and structured tool calling. Models are available via API, cloud hosters, and open-weights download repositories.
Model Lineage & Specification
deepseek-v4-flash is deepseek-ai's primary release in the current family.
deepseek-v4-flashComparable models
| Feature | DeepSeek-V4-Flash | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Flash-0731-GGUF | Claude Opus 5 |
|---|---|---|---|---|
| Description | DeepSeek-V4-Flash is a 284B-parameter Mixture-of-Experts (MoE) text generation m... | DeepSeek-V4-Flash-0731 is a 304.2B-parameter Mixture-of-Experts (MoE) text gener... | The DeepSeek-V4 series introduces next-generation Mixture-of-Experts (MoE) open ... | For complex agentic coding and enterprise work. Claude Opus 5 is a step-change i... |
| API Identifier | deepseek-v4-flash | deepseek-v4-flash-0731 | deepseek-v4-flash-0731-gguf | anthropic-claude-opus-5 |
| Parameters | 284 Billion | 304.2B | 284B parameters | — |
| Context Window | 1M | unknown | 1M | 1M tokens |
| License / Type | open weights | open weights | open weights | api only |
DeepSeek-V4-Flash is a 284B-parameter Mixture-of-Experts (MoE) text generation model developed by deepseek-ai. Built on the DeepseekV4ForCausalLM architecture using transformers, it activates 13B parameters per token. Released on 2026-04-22 with 1,945 likes and 2,814,414 downloads on Hugging Face (where compressed file configurations cause systems to display it as 158.1B).
DeepSeek-V4-Flash
Model Overview
DeepSeek-V4-Flash is a 158.1B-parameter Mixture-of-Experts text generation model developed by deepseek-ai. Built on the DeepseekV4ForCausalLM architecture. Released on 2026-04-22.
📊 Quick Specs
| Specification | Value |
|---|---|
| Parameters | 158.1B |
| Architecture | DeepseekV4ForCausalLM |
| Task | text generation |
| Modality | text |
| License | MIT |
| Framework | transformers |
| MoE | Yes |
| Languages | — |
✨ Key Features
- 158.1B parameters with sparse MoE architecture for efficient inference
- Built on DeepseekV4ForCausalLM architecture (transformers)
- Primary task: text generation (text modality)
- Open-weights under MIT license — self-hostable and fine-tunable
- High adoption: 2,814,414 downloads on Hugging Face
📈 Community Adoption
- 1,945 likes on Hugging Face
- 2,814,414 downloads on Hugging Face
🔗 Resources
- Hugging Face Hub: DeepSeek-V4-Flash on Hugging Face
- Paper: arXiv
📜 License & Access
MIT — Open-weights model available for download, fine-tuning, and self-hosted deployment.