Compare Models
Evaluate leading AI models side-by-side. Compare context windows, open-weights licensing, parameters, and verified benchmarks to choose the right model for your application.
Comparing 1 of 4 models
Model Comparison Matrix
Add a model | Add a model | Add a model | ||
|---|---|---|---|---|
| Type | API Only | |||
| Release Date | Jun 15, 2026 | |||
| License | Proprietary | |||
| Parameters | undisclosed | |||
| Modalities | textimage | |||
| Context Window | 1M tokens | |||
| Primary Task | chat reasoning | |||
| Release Date | 2026-06-15 | |||
| CTI-REALM | 67.5 (score) | |||
| CharXiv Reasoning | 85.1 (score) | |||
| GPQA Diamond | 95.5 (score) | |||
| Humanity’s Last Exam | 47.2 (score) | |||
| LiveCodeBench | 92.9 (score) | |||
| LiveCodeBench Pro | 87.8 (score) | |||
| Long Context Reasoning | 74.7 (score) | |||
| MRCRv2 | 86.6 (score) | |||
| SWE Bench Pro | 59 (score) | |||
| SciCode | 60.1 (score) | |||
| Terminal Bench 2.1 | 80.2 (score) | |||
| τ3 Banking | 21.7 (score) |