Science benchmarks

Quantitative science and reasoning performance on graduate-level evaluations. Every score is sourced and citable — see the provenance notes below before comparing across rows or variants.

Model
GPT-5.5
openai/gpt-5.5
43.6% 📊
2026-07
Claude Sonnet 5
anthropic/claude-sonnet-5
43.2% 🏢
2026-06
57.4% 🏢
2026-06
Claude Opus 4.8
anthropic/claude-opus-4-8
42.2% 📊
2026-07
57.9% 🏢
2026-06
GPT-5.4 mini
openai/gpt-5.4-mini
23.5% 📊
2026-07
Gemini 2.5 Pro
google-deepmind/gemini-2.5-pro
21.6% 📊
2026-07
3.8% 🔬
2025-05
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
21.1% 📊
2026-07
46.8% 🏢
2026-06
o3
openai/o3
20.3% 📊
2026-07
6.5% 🔬
2025-05
Gemini 2.5 Flash
google-deepmind/gemini-2.5-flash
12.1% 📊
2026-07
Claude Sonnet 4
anthropic/claude-sonnet-4
9.6% 📊
2026-07
5.9% 🔬
2025-05
Claude 3.5 Sonnet
anthropic/claude-3-5-sonnet
8.9% 📄
2025-01
DeepSeek R1
deepseek/deepseek-r1
8.5% 📊
2025-04
1.3% 🔬
2025-05
GPT-4o
openai/gpt-4o
2.7% 📊
2026-07

Sourcing and comparability

Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.

HLE variant warning: HLE (no tools) and HLE (with tools) are separate evaluations on the same question corpus and must not be compared directly. With-tools scores are systematically 10–20 percentage points higher than no-tools scores for the same model. GPT-4o and Claude 3.5 Sonnet entries show only no-tools scores from the original January 2025 benchmark paper — they are historical baselines, not current performance figures.

Vendor vs. leaderboard gap: Some models have both a vendor-reported score and a CAIS leaderboard score for the same benchmark. These differ because vendor announcements often use extended thinking, optimised prompting, or internal test sets, while the CAIS AI Dashboard uses standardised conditions across all models. The table shows the most recent entry per benchmark — check the source badge and click the ↗ link to see which evaluation applies. For Anthropic models the typical gap is 7–14 percentage points.

ARC-AGI-2 compute sensitivity: ARC-AGI-2 scores are far more compute-sensitive than HLE. The same model can score 2–4× higher with extended thinking tokens or a higher-compute configuration vs. the default API. The scores shown here are from the ARC Prize Foundation's "We tested every major AI reasoning system" evaluation at ARC-AGI-2's May 2025 launch, which used matched compute settings across models. Only 4 of the 12 tracked models were included in that evaluation — the remaining cells show "—". Human average on ARC-AGI-2 is approximately 66%.

Data is growing: This table is sourced from modelglass-science, a registry of science and reasoning benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.