Compare Models

Evaluate leading AI models side-by-side. Compare context windows, open-weights licensing, parameters, and verified benchmarks to choose the right model for your application.

Comparing 1 of 4 models
Model Comparison Matrix
Add a model
Add a model
Add a model
TypeAPI Only
Release DateJun 15, 2026
LicenseProprietary
Parametersundisclosed
Modalities
textimage
Context Window1M tokens
Primary Taskchat reasoning
Release Date2026-06-15
CTI-REALM
67.5 (score)
CharXiv Reasoning
85.1 (score)
GPQA Diamond
95.5 (score)
Humanity’s Last Exam
47.2 (score)
LiveCodeBench
92.9 (score)
LiveCodeBench Pro
87.8 (score)
Long Context Reasoning
74.7 (score)
MRCRv2
86.6 (score)
SWE Bench Pro
59 (score)
SciCode
60.1 (score)
Terminal Bench 2.1
80.2 (score)
τ3 Banking
21.7 (score)