AI Model Benchmarks

18 source-dated benchmark snapshots, curated by AI Explained, author of SimpleBench

Evaluation settings and metrics vary; use each card's linked source before comparing scores.

Compare Models

Humanity's Last Exam

Results snapshot Jun 14, 2026 • Source checked Aug 18, 2026

2,500 of the toughest, subject-diverse, multi-modal questions designed to test for both depth of reasoning and breadth of knowledge. Created in partnership with the Center for AI Safety, HLE includes questions across mathematics, humanities, and natural sciences from nearly 1,000 expert contributors.
Models (no tools)Score
1Gemini 3.1 Pro Preview (high thinking)46.44% ±1.96
1GPT-5.4 Pro44.32% ±1.95
3Muse Spark40.56% ±1.92
3Gemini 3 Pro Preview37.52% ±1.90
4GPT-5.4 (xhigh)36.24% ±1.88
Full Results

SimpleBench

Results snapshot Aug 13, 2026 • Source checked Aug 18, 2026

Asks "trick" questions that require common-sense reasoning rather than memorized facts. Models must avoid being misled by common traps.
ModelScore
1Claude Fable 581.9%
2Claude Opus 580.6%
3Gemini 3.1 Pro Preview79.6%
4GPT-5.5 Pro76.9%
5Gemini 3.5 Flash76.7%
Full Results

METR Time Horizons

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

METR's time horizon is the human task duration at which an AI model reaches 50% success. Tasks are drawn from RE-Bench, HCAST and SWAA, which cover machine-learning research engineering, general software engineering and software operations respectively.
ModelMinutes
1Claude Mythos Preview (early)1044.8 ±1397.7
2Claude Opus 4.6 (unknown settings)718.8 ±1815.2
3Gemini 3.1 Pro Preview384.1 ±230.6
4GPT-5.2 (high)352.2 ±335.5
5GPT-5.3 Codex349.5 ±333.1
Full Results

SWE-bench Verified

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

A human-curated subset of 500 GitHub issues from the SWE-bench dataset tests whether models can implement valid code fixes. The model interacts with a Python repository and must modify the correct files to fix the issue. The solution is judged by running unit tests.
ModelScore
1Claude Opus 4.7 (max)83.5% ±1.7
2GPT-5.5 (xhigh)80.6% ±1.8
3Gemini 3.5 Flash (high)79.3% ±1.8
4Claude Opus 4.6 (no thinking)78.7% ±1.9
5GLM 5.2 (max)78.7% ±1.9
Full Results

GPQA Diamond

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

A multiple-choice set of 198 PhD-level science questions in biology, chemistry and physics. It focuses on "Diamond" items for which domain experts answered correctly while non-experts often failed; random guessing yields ~25%.
ModelScore
1GPT-5.4 Pro (xhigh)94.6% ±1.6
2Gemini 3.1 Pro Preview94.1% ±1.7
3GPT-5.5 (xhigh)94.0% ±1.5
4GPT-5.5 Pro (xhigh)93.9% ±1.6
5GPT-5.4 (xhigh)93.3% ±1.8
Full Results

GDPval

Results snapshot Dec 11, 2025 • Source checked Aug 18, 2026

GDPval is an OpenAI-led benchmark spanning 44 knowledge-work occupations across nine major U.S. industries. This historical snapshot reports clear-win rates against work produced by experienced professionals; OpenAI has retired the original hosted leaderboard.
ModelScore
1GPT-5.249.8%
2Claude Opus 4.545.5%
3Claude Opus 4.143.6%
4Claude Sonnet 4.542.5%
5Gemini 3 Pro Preview40.3%
Full Results

Text Arena (Coding)

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

Previously known as WebDev Arena, this benchmark pits models against each other to build websites or web apps from prompts. Voters choose the better result, and scores are calculated with a Bradley-Terry model.
ModelScore
1Claude Fable 5 (unknown)1653.9
2GLM 5.2 (max)1593.3
3Claude Opus 4.7 (unknown settings)1566.8
4Claude Opus 4.7 (no thinking)1562.4
5Claude Opus 4.6 (unknown settings)1556.3
Full Results

GSO (General Speedup Optimization)

Results snapshot Jul 12, 2026 • Source checked Aug 18, 2026

Assesses models' ability to optimize software performance. Each task requires making code changes within a limited number of attempts. Performance is measured using OPT@K - the percentage of tasks where the model achieves at least 95% of the human-achieved speedup.
ModelScore
1Claude Opus 4.8 (xhigh)47.1%
2Claude Opus 4.7 (high)44.1%
3Claude Opus 4.6 (high)41.2%
4GPT-5.5 (xhigh)40.2%
5Claude Sonnet 5 (xhigh)37.3%
Full Results

Fiction.liveBench

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

Models answer questions about long serialized stories hosted on Fiction.live. The card uses the 16k-token score, testing recall, chronology, character understanding and inference.
ModelScore
1GPT-5 (medium)97.2%
2o3-pro97.2%
3Grok 494.4%
4Grok 4 Fast94.4%
5Gemini 2.5 Pro Preview (Jun '25)91.7%
Full Results

BALROG

Results snapshot Feb 25, 2026 • Source checked Aug 18, 2026

Evaluates agents across BabyAI, Crafter, TextWorld, Baba Is AI, NetHack and MiniHack. The progress score averages task completion across repeated seeded runs.
ModelScore
1Gemini 3 Pro Preview58.1% ±2.1
2Gemini 3.1 Pro Preview57.0% ±2.0
3Gemini 3 Flash48.1% ±2.4
4Grok 443.6% ±2.2
5Claude Opus 4.543.5% ±2.3
Full Results

OTIS Mock AIME 2024-25

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

This benchmark uses 45 integer-answer problems from unofficial Mock AIME exams (2024-2025). Problems are harder than MATH Level 5 but easier than FrontierMath and have answers between 0 and 999.
ModelScore
1GPT-5.5 Pro (xhigh)100.0% ±0.0
2GPT-5.5 (xhigh)100.0% ±0.0
3Claude Fable 5 (max)99.7% ±0.3
4Claude Opus 4.898.3% ±1.4
5Claude Opus 4.7 (xhigh)97.8% ±2.2
Full Results

MATH Level 5

Results snapshot Oct 30, 2025 • Source checked Aug 18, 2026

The Level 5 subset of the MATH dataset contains the hardest competition-style problems from AMC 10, AMC 12 and AIME. Answers are scored using a combination of normalized string match, symbolic equivalence and model-graded equivalence.
ModelScore
1GPT-5 (high)98.1% ±0.3
2GPT-5 (medium)97.9% ±0.3
3GPT-5 mini (high)97.8% ±0.3
4o4-mini (high)97.8% ±0.3
5o3 (high)97.8% ±0.3
Full Results

FrontierMath Tiers 1-3 (v2)

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

Expert-written unpublished mathematics problems covering advanced undergraduate through early-career research difficulty. Epoch released v2 on June 12, 2026 after correcting and removing problematic items.
ModelScore
1GPT-5.5 Pro (xhigh)87.7% ±1.9
2Claude Fable 5 (max)87.0% ±2.0
3GPT-5.5 (xhigh)85.3% ±2.1
4GPT-5.4 Pro (xhigh)82.5% ±2.3
5Claude Opus 4.8 (max)80.0% ±2.4
Full Results

FrontierMath Tier 4 (v2)

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

Research-level FrontierMath problems from Epoch AI. Tier 4 contains the hardest private problems in the v2 FrontierMath release.
ModelScore
1Claude Fable 5 (max)87.8% ±5.2
2GPT-5.5 Pro (xhigh)78.0% ±6.5
3AI co-mathematician75.6% ±6.7
4GPT-5.5 (xhigh)72.5% ±7.1
5GPT-5.4 Pro (xhigh)58.5% ±7.8
Full Results

WeirdML v2

Results snapshot Aug 17, 2026 • Source checked Aug 18, 2026

Asks models to write code that trains machine-learning models to solve non-standard tasks (e.g., recognizing shapes, classifying digits, predicting chess outcomes). Models iterate on code, training and evaluating within a constrained environment.
ModelScore
1Claude Fable 5 (max)91.9%
2Claude Opus 5 (max)91.8%
3Claude Opus 5 (high)91.6%
4GPT-5.6 Sol Pro (max)89.4%
5GPT-5.6 Sol (high)88.8%
Full Results

Terminal-Bench 2.0

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

Models use a terminal to complete assignments such as editing files, running commands and debugging code. The tasks come from various agent frameworks; performance is the success rate.
ModelScore
1Claude Opus 4.7 (unknown settings)90.2%
2GPT-5.5 (unknown thinking)84.7%
3GPT-5.4 (none)81.8%
4Gemini 3.1 Pro Preview80.2%
5Claude Opus 4.6 (unknown settings)79.8%
Full Results

VPCT (Visual Physics Comprehension Test)

Results snapshot Mar 6, 2026 • Source checked Aug 18, 2026

Each problem shows an image of a ramp with buckets; the model must predict in which bucket a ball will land. Tasks test basic understanding of gravity and motion.
ModelScore
1Gemini 3 Pro Preview91.0%
2GPT-5.2 (xhigh)84.0%
3Gemini 3 Flash72.6%
4GPT-5.2 (high)67.0%
5GPT-5 (high)66.0%
Full Results

GeoBench

Results snapshot Aug 18, 2026 • Source checked Aug 18, 2026

Inspired by GeoGuessr. Models inspect street-level photos and guess the country and coordinates. The card uses country accuracy on the A Common World map; the benchmark also reports distance-based scores.
ModelScore
1Gemini 3 Flash88.0%
2Gemini 2.5 Pro Preview (Jun '25)86.0%
3Gemini 3 Pro Preview84.0%
4Gemini 2.5 Pro Exp (Mar '25)81.0%
5GPT-5 (medium)81.0%
Full Results

Newest results snapshot August 18, 2026 • Each benchmark links its source and checked date • Epoch AI data: Epoch AI (CC BY)