Compare Models

Evaluate leading AI models side-by-side. Compare context windows, open-weights licensing, parameters, and verified benchmarks to choose the right model for your application.

Comparing 1 of 4 models

Compare Agent & Coding Benchmarks

Visualize side-by-side relative performance differences for selected models on core coding benchmarks.

DeepSeek-V4-Flash-0731-GGUF

Dynamic Running Cost Estimator

Estimate query costs based on the model pricing database and your expected volume.

DeepSeek-V4-Flash-0731-GGUFFree (Self-Hosted)
Not all models have been independently verified yet. Benchmark scores display exact field confidence (VERIFIED / LIKELY / DRAFT). Uncorroborated benchmarks show Pending Curator Audit.
Model Comparison Matrix
Add a model
Add a model
Add a model
TypeOpen Weights
LicenseMIT
Parameters284B parameters
Modalities
text
Context Window1MLargest
Estimated cost / queryFreeCheapest
SWE-Bench Score(Pending Audit)
Aider Polyglot(Pending Audit)
GPQA Diamond(Pending Audit)
Agents' Last Exam
25.2%Verified
AutomationBench
25.1%Verified
Cybergym
76.7%Verified
DSBench-Hard t
59.6%Verified
DeepSWE
54.4%Verified
NL2Repo
54.2%Verified
Public DSBench-FullStack †
68.7%Verified
Terminal Bench 2.1
82.7%Verified
Toolathlon-Verified
70.3%Verified