Models & pricing / Models / DeepSeek-V4-Flash

DeepSeek-V4-Flash Overview

DeepSeek-V4-Flash is a 284B-parameter Mixture-of-Experts (MoE) text generation model developed by deepseek-ai. Built on the DeepseekV4ForCausalLM architecture using transformers, it activates 13B parameters per token. Released on 2026-04-22 with 1,945 likes and 2,814,414 downloads on Hugging Face (where compressed file configurations cause systems to display it as 158.1B).

Choosing a model

If you're unsure which model to use, start with DeepSeek-V4-Flash for complex agentic coding, reasoning, and enterprise workloads. For lightweight, low-latency autocomplete or edge tasks, consider smaller parameters.

All current deepseek-ai models support text and multimodal input, multilingual reasoning, and structured tool calling. Models are available via API, cloud hosters, and open-weights download repositories.

Model Lineage & Specification

deepseek-v4-flash is deepseek-ai's primary release in the current family.

Developerdeepseek-ai
API Identifierdeepseek-v4-flash
Parameters284 Billion
Context Window1M
LicenseMIT

Comparable models

FeatureDeepSeek-V4-FlashDeepSeek-V4-Flash-0731DeepSeek-V4-Flash-0731-GGUFClaude Opus 5
DescriptionDeepSeek-V4-Flash is a 284B-parameter Mixture-of-Experts (MoE) text generation m...DeepSeek-V4-Flash-0731 is a 304.2B-parameter Mixture-of-Experts (MoE) text gener...The DeepSeek-V4 series introduces next-generation Mixture-of-Experts (MoE) open ...For complex agentic coding and enterprise work. Claude Opus 5 is a step-change i...
API Identifierdeepseek-v4-flashdeepseek-v4-flash-0731deepseek-v4-flash-0731-ggufanthropic-claude-opus-5
Parameters284 Billion304.2B284B parameters
Context Window1Munknown1M1M tokens
License / Typeopen weightsopen weightsopen weightsapi only

DeepSeek-V4-Flash is a 284B-parameter Mixture-of-Experts (MoE) text generation model developed by deepseek-ai. Built on the DeepseekV4ForCausalLM architecture using transformers, it activates 13B parameters per token. Released on 2026-04-22 with 1,945 likes and 2,814,414 downloads on Hugging Face (where compressed file configurations cause systems to display it as 158.1B).

DeepSeek-V4-Flash

Model Overview

DeepSeek-V4-Flash is a 158.1B-parameter Mixture-of-Experts text generation model developed by deepseek-ai. Built on the DeepseekV4ForCausalLM architecture. Released on 2026-04-22.


📊 Quick Specs

SpecificationValue
Parameters158.1B
ArchitectureDeepseekV4ForCausalLM
Tasktext generation
Modalitytext
LicenseMIT
Frameworktransformers
MoEYes
Languages

✨ Key Features

  • 158.1B parameters with sparse MoE architecture for efficient inference
  • Built on DeepseekV4ForCausalLM architecture (transformers)
  • Primary task: text generation (text modality)
  • Open-weights under MIT license — self-hostable and fine-tunable
  • High adoption: 2,814,414 downloads on Hugging Face

📈 Community Adoption

  • 1,945 likes on Hugging Face
  • 2,814,414 downloads on Hugging Face

🔗 Resources


📜 License & Access

MIT — Open-weights model available for download, fine-tuning, and self-hosted deployment.