Compare Models

Evaluate leading AI models side-by-side. Compare context windows, open-weights licensing, parameters, and verified benchmarks to choose the right model for your application.

Comparing 1 of 4 models
Model Comparison Matrix
Add a model
Add a model
Add a model
TypeAPI Only
Release DateJun 15, 2026
LicenseProprietary
Parametersundisclosed
Modalities
textimage
Context Window1M tokens
Primary Taskchat reasoning
Release Date2026-06-15
CTI-REALM
69.4 (score)
CharXiv Reasoning
86.6 (score)
GPQA Diamond
95.5 (score)
Humanity’s Last Exam
50 (score)
LiveCodeBench
93.2 (score)
LiveCodeBench Pro
90.8 (score)
Long Context Reasoning
73.3 (score)
MRCRv2
93.6 (score)
SWE Bench Pro
73.7 (score)
SciCode
58.7 (score)
Terminal Bench 2.1
82.1 (score)
τ3 Banking
20.6 (score)