Compare Models

Evaluate leading AI models side-by-side. Compare context windows, open-weights licensing, parameters, and verified benchmarks to choose the right model for your application.

Comparing 1 of 4 models
Model Comparison Matrix
Add a model
Add a model
Add a model
TypeAPI Only
Release DateMay 21, 2026
LicenseProprietary
Parametersundisclosed
Modalities
text
Context Window1M tokens
Primary Taskchat reasoning
Release Date2026-05-21
GPQA Diamond
92.4 (accuracy)
Humanity's Last Exam
41.4 (accuracy)
MCP Atlas
76.4%
NL2Repo
47.2 (score)
SWE-Bench Multilingual
78.3 (resolve rate)
SWE-Bench Pro
60.6 (resolve rate)
SWE-Bench Verified
80.4%
SciCode
53.5 (score)
Terminal-Bench
69.7%