Science benchmarks
Quantitative science and reasoning performance on graduate-level evaluations. Every score is sourced and citable — see the provenance notes below before comparing across rows or variants.
| Model | |||
|---|---|---|---|
| GPT-5.5 openai/gpt-5.5 | 43.6% 📊 ↗ 2026-07 | — | — |
| Claude Sonnet 5 anthropic/claude-sonnet-5 | 43.2% 🏢 ↗ 2026-06 | 57.4% 🏢 ↗ 2026-06 | — |
| Claude Opus 4.8 anthropic/claude-opus-4-8 | 42.2% 📊 ↗ 2026-07 | 57.9% 🏢 ↗ 2026-06 | — |
| GPT-5.4 mini openai/gpt-5.4-mini | 23.5% 📊 ↗ 2026-07 | — | — |
| Gemini 2.5 Pro google-deepmind/gemini-2.5-pro | 21.6% 📊 ↗ 2026-07 | — | 3.8% 🔬 ↗ 2025-05 |
| Claude Sonnet 4.6 anthropic/claude-sonnet-4-6 | 21.1% 📊 ↗ 2026-07 | 46.8% 🏢 ↗ 2026-06 | — |
| o3 openai/o3 | 20.3% 📊 ↗ 2026-07 | — | 6.5% 🔬 ↗ 2025-05 |
| Gemini 2.5 Flash google-deepmind/gemini-2.5-flash | 12.1% 📊 ↗ 2026-07 | — | — |
| Claude Sonnet 4 anthropic/claude-sonnet-4 | 9.6% 📊 ↗ 2026-07 | — | 5.9% 🔬 ↗ 2025-05 |
| Claude 3.5 Sonnet anthropic/claude-3-5-sonnet | 8.9% 📄 ↗ 2025-01 | — | — |
| DeepSeek R1 deepseek/deepseek-r1 | 8.5% 📊 ↗ 2025-04 | — | 1.3% 🔬 ↗ 2025-05 |
| GPT-4o openai/gpt-4o | 2.7% 📊 ↗ 2026-07 | — | — |
Sourcing and comparability
Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.
HLE variant warning: HLE (no tools) and HLE (with tools) are separate evaluations on the same question corpus and must not be compared directly. With-tools scores are systematically 10–20 percentage points higher than no-tools scores for the same model. GPT-4o and Claude 3.5 Sonnet entries show only no-tools scores from the original January 2025 benchmark paper — they are historical baselines, not current performance figures.
Vendor vs. leaderboard gap: Some models have both a vendor-reported score and a CAIS leaderboard score for the same benchmark. These differ because vendor announcements often use extended thinking, optimised prompting, or internal test sets, while the CAIS AI Dashboard uses standardised conditions across all models. The table shows the most recent entry per benchmark — check the source badge and click the ↗ link to see which evaluation applies. For Anthropic models the typical gap is 7–14 percentage points.
ARC-AGI-2 compute sensitivity: ARC-AGI-2 scores are far more compute-sensitive than HLE. The same model can score 2–4× higher with extended thinking tokens or a higher-compute configuration vs. the default API. The scores shown here are from the ARC Prize Foundation's "We tested every major AI reasoning system" evaluation at ARC-AGI-2's May 2025 launch, which used matched compute settings across models. Only 4 of the 12 tracked models were included in that evaluation — the remaining cells show "—". Human average on ARC-AGI-2 is approximately 66%.
Data is growing: This table is sourced from modelglass-science, a registry of science and reasoning benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.