18 source-dated benchmark snapshots, curated by AI Explained, author of SimpleBench
Evaluation settings and metrics vary; use each card's linked source before comparing scores.
Results snapshot Jun 14, 2026 • Source checked Aug 18, 2026
| Models (no tools) | Score | |
|---|---|---|
| 1 | Gemini 3.1 Pro Preview (high thinking) | 46.44% ±1.96 |
| 1 | GPT-5.4 Pro | 44.32% ±1.95 |
| 3 | Muse Spark | 40.56% ±1.92 |
| 3 | Gemini 3 Pro Preview | 37.52% ±1.90 |
| 4 | GPT-5.4 (xhigh) | 36.24% ±1.88 |
Results snapshot Aug 13, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Claude Fable 5 | 81.9% |
| 2 | Claude Opus 5 | 80.6% |
| 3 | Gemini 3.1 Pro Preview | 79.6% |
| 4 | GPT-5.5 Pro | 76.9% |
| 5 | Gemini 3.5 Flash | 76.7% |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Minutes | |
|---|---|---|
| 1 | Claude Mythos Preview (early) | 1044.8 ±1397.7 |
| 2 | Claude Opus 4.6 (unknown settings) | 718.8 ±1815.2 |
| 3 | Gemini 3.1 Pro Preview | 384.1 ±230.6 |
| 4 | GPT-5.2 (high) | 352.2 ±335.5 |
| 5 | GPT-5.3 Codex | 349.5 ±333.1 |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Claude Opus 4.7 (max) | 83.5% ±1.7 |
| 2 | GPT-5.5 (xhigh) | 80.6% ±1.8 |
| 3 | Gemini 3.5 Flash (high) | 79.3% ±1.8 |
| 4 | Claude Opus 4.6 (no thinking) | 78.7% ±1.9 |
| 5 | GLM 5.2 (max) | 78.7% ±1.9 |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | GPT-5.4 Pro (xhigh) | 94.6% ±1.6 |
| 2 | Gemini 3.1 Pro Preview | 94.1% ±1.7 |
| 3 | GPT-5.5 (xhigh) | 94.0% ±1.5 |
| 4 | GPT-5.5 Pro (xhigh) | 93.9% ±1.6 |
| 5 | GPT-5.4 (xhigh) | 93.3% ±1.8 |
Results snapshot Dec 11, 2025 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | GPT-5.2 | 49.8% |
| 2 | Claude Opus 4.5 | 45.5% |
| 3 | Claude Opus 4.1 | 43.6% |
| 4 | Claude Sonnet 4.5 | 42.5% |
| 5 | Gemini 3 Pro Preview | 40.3% |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Claude Fable 5 (unknown) | 1653.9 |
| 2 | GLM 5.2 (max) | 1593.3 |
| 3 | Claude Opus 4.7 (unknown settings) | 1566.8 |
| 4 | Claude Opus 4.7 (no thinking) | 1562.4 |
| 5 | Claude Opus 4.6 (unknown settings) | 1556.3 |
Results snapshot Jul 12, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Claude Opus 4.8 (xhigh) | 47.1% |
| 2 | Claude Opus 4.7 (high) | 44.1% |
| 3 | Claude Opus 4.6 (high) | 41.2% |
| 4 | GPT-5.5 (xhigh) | 40.2% |
| 5 | Claude Sonnet 5 (xhigh) | 37.3% |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | GPT-5 (medium) | 97.2% |
| 2 | o3-pro | 97.2% |
| 3 | Grok 4 | 94.4% |
| 4 | Grok 4 Fast | 94.4% |
| 5 | Gemini 2.5 Pro Preview (Jun '25) | 91.7% |
Results snapshot Feb 25, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Gemini 3 Pro Preview | 58.1% ±2.1 |
| 2 | Gemini 3.1 Pro Preview | 57.0% ±2.0 |
| 3 | Gemini 3 Flash | 48.1% ±2.4 |
| 4 | Grok 4 | 43.6% ±2.2 |
| 5 | Claude Opus 4.5 | 43.5% ±2.3 |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | GPT-5.5 Pro (xhigh) | 100.0% ±0.0 |
| 2 | GPT-5.5 (xhigh) | 100.0% ±0.0 |
| 3 | Claude Fable 5 (max) | 99.7% ±0.3 |
| 4 | Claude Opus 4.8 | 98.3% ±1.4 |
| 5 | Claude Opus 4.7 (xhigh) | 97.8% ±2.2 |
Results snapshot Oct 30, 2025 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | GPT-5 (high) | 98.1% ±0.3 |
| 2 | GPT-5 (medium) | 97.9% ±0.3 |
| 3 | GPT-5 mini (high) | 97.8% ±0.3 |
| 4 | o4-mini (high) | 97.8% ±0.3 |
| 5 | o3 (high) | 97.8% ±0.3 |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | GPT-5.5 Pro (xhigh) | 87.7% ±1.9 |
| 2 | Claude Fable 5 (max) | 87.0% ±2.0 |
| 3 | GPT-5.5 (xhigh) | 85.3% ±2.1 |
| 4 | GPT-5.4 Pro (xhigh) | 82.5% ±2.3 |
| 5 | Claude Opus 4.8 (max) | 80.0% ±2.4 |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Claude Fable 5 (max) | 87.8% ±5.2 |
| 2 | GPT-5.5 Pro (xhigh) | 78.0% ±6.5 |
| 3 | AI co-mathematician | 75.6% ±6.7 |
| 4 | GPT-5.5 (xhigh) | 72.5% ±7.1 |
| 5 | GPT-5.4 Pro (xhigh) | 58.5% ±7.8 |
Results snapshot Aug 17, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Claude Fable 5 (max) | 91.9% |
| 2 | Claude Opus 5 (max) | 91.8% |
| 3 | Claude Opus 5 (high) | 91.6% |
| 4 | GPT-5.6 Sol Pro (max) | 89.4% |
| 5 | GPT-5.6 Sol (high) | 88.8% |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Claude Opus 4.7 (unknown settings) | 90.2% |
| 2 | GPT-5.5 (unknown thinking) | 84.7% |
| 3 | GPT-5.4 (none) | 81.8% |
| 4 | Gemini 3.1 Pro Preview | 80.2% |
| 5 | Claude Opus 4.6 (unknown settings) | 79.8% |
Results snapshot Mar 6, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Gemini 3 Pro Preview | 91.0% |
| 2 | GPT-5.2 (xhigh) | 84.0% |
| 3 | Gemini 3 Flash | 72.6% |
| 4 | GPT-5.2 (high) | 67.0% |
| 5 | GPT-5 (high) | 66.0% |
Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026
| Model | Score | |
|---|---|---|
| 1 | Gemini 3 Flash | 88.0% |
| 2 | Gemini 2.5 Pro Preview (Jun '25) | 86.0% |
| 3 | Gemini 3 Pro Preview | 84.0% |
| 4 | Gemini 2.5 Pro Exp (Mar '25) | 81.0% |
| 5 | GPT-5 (medium) | 81.0% |