Science benchmarks

Quantitative science and reasoning performance on graduate-level evaluations. Every score is sourced and citable — see the provenance notes below before comparing across rows or variants.

Frequently asked questions

What is HLE (no tools)?
Humanity's Last Exam (HLE) is a 2,500-question benchmark spanning dozens of academic subjects — mathematics, humanities, and the natural sciences — created because earlier benchmarks like MMLU had become saturated (frontier models were scoring over 90%). Each question has a single, unambiguous, expert-verified answer that can't be found via a quick web search; 80% are exact-match, the rest multiple-choice. This page tracks two variants on the same question set — 'no tools' (closed-book, the anchor benchmark here) and 'with tools' (web search, calculator, and code execution enabled) — plus ARC-AGI-2, a separate abstract visual-reasoning benchmark. See the sourcing notes below for how the variants and ARC-AGI-2 compare.
Which models are currently tracked?
14 models: Claude Opus 5, GPT-5.5, Claude Sonnet 5, Claude Opus 4.8, GPT-5.4 mini, Gemini 2.5 Pro, Claude Sonnet 4.6, o3, Gemini 2.5 Flash, Claude Sonnet 4, Claude 3.5 Sonnet, DeepSeek R1, GPT-4o, Gemini 3.5 Flash.
Which model scores highest on HLE (no tools)?
Claude Opus 5, at 56.3% as of 2026-07.
How often is the data updated?
This page was last rebuilt August 10, 2026 from the modelglass-science registry. Scores are appended as new evaluations are verified — an existing result is never edited or deleted, only superseded by a newer entry.
Where does the score come from?
Every score below links to its source (the ↗ icon next to each result) and shows an icon for how it was obtained: 🏢 vendor-reported (the model developer's own announcement), 📊 public leaderboard (a public, independently-maintained evaluation run — for HLE, typically the CAIS AI Safety Dashboard), 🔬 independent evaluation (a credible third party ran the benchmark), or 📄 published paper (a peer-reviewed or arXiv result). See the sourcing notes below the table for the HLE tool-variant comparison, the vendor-vs-leaderboard score gap, and ARC-AGI-2's compute sensitivity.

Who this is for

This page is built for researchers, model evaluators, and technically-minded readers who want a domain-agnostic read on frontier reasoning and knowledge capability, not a specific task outcome. Humanity's Last Exam was created explicitly to give researchers and policymakers a benchmark that hasn't saturated the way MMLU and similar exams have; ARC-AGI-2 serves a related but distinct audience — those studying whether a model reasons and generalizes rather than pattern-matches its way to an answer.

Model
Claude Opus 5
anthropic/claude-opus-5
56.3% 🏢
2026-07
64.7% 🏢
2026-07
90.4% 🔬
2026-07
GPT-5.5
openai/gpt-5.5
43.6% 📊
2026-07
Claude Sonnet 5
anthropic/claude-sonnet-5
43.2% 🏢
2026-06
57.4% 🏢
2026-06
Claude Opus 4.8
anthropic/claude-opus-4-8
42.2% 📊
2026-07
57.9% 🏢
2026-06
GPT-5.4 mini
openai/gpt-5.4-mini
23.5% 📊
2026-07
Gemini 2.5 Pro
google-deepmind/gemini-2.5-pro
21.6% 📊
2026-07
3.8% 🔬
2025-05
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
21.1% 📊
2026-07
46.8% 🏢
2026-06
o3
openai/o3
20.3% 📊
2026-07
6.5% 🔬
2025-05
Gemini 2.5 Flash
google-deepmind/gemini-2.5-flash
12.1% 📊
2026-07
Claude Sonnet 4
anthropic/claude-sonnet-4
9.6% 📊
2026-07
5.9% 🔬
2025-05
Claude 3.5 Sonnet
anthropic/claude-3-5-sonnet
8.9% 📄
2025-01
DeepSeek R1
deepseek/deepseek-r1
8.5% 📊
2025-04
1.3% 🔬
2025-05
GPT-4o
openai/gpt-4o
2.7% 📊
2026-07
Gemini 3.5 Flash
google-deepmind/gemini-3.5-flash
72.1% 🔬
2026-05

See how benchmark data like this can power a cost-aware routing setup

See how →

Why we track this

We track HLE (no tools) because it's a genuine differentiator between models — not a vanity metric. Used alongside the other benchmarks on this page and pricing data elsewhere in the registry, it's designed to help you find the most appropriate model for your specific use case, not just the highest-scoring one.

Sourcing and comparability

Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.

HLE variant warning: HLE (no tools) and HLE (with tools) are separate evaluations on the same question corpus and must not be compared directly. With-tools scores are systematically 10–20 percentage points higher than no-tools scores for the same model. GPT-4o and Claude 3.5 Sonnet entries show only no-tools scores from the original January 2025 benchmark paper — they are historical baselines, not current performance figures.

Vendor vs. leaderboard gap: Some models have both a vendor-reported score and a CAIS leaderboard score for the same benchmark. These differ because vendor announcements often use extended thinking, optimised prompting, or internal test sets, while the CAIS AI Dashboard uses standardised conditions across all models. The table shows the most recent entry per benchmark — check the source badge and click the ↗ link to see which evaluation applies. For Anthropic models the typical gap is 7–14 percentage points.

ARC-AGI-2 compute sensitivity: ARC-AGI-2 scores are far more compute-sensitive than HLE. The same model can score 2–4× higher with extended thinking tokens or a higher-compute configuration vs. the default API. The scores shown here are from the ARC Prize Foundation's "We tested every major AI reasoning system" evaluation at ARC-AGI-2's May 2025 launch, which used matched compute settings across models. Only 4 of the 12 tracked models were included in that evaluation — the remaining cells show "—". Human average on ARC-AGI-2 is approximately 66%.

Data is growing: This table is sourced from modelglass-science, a registry of science and reasoning benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.

Related pages