Agentic benchmarks
Quantitative agentic performance on real-world, tool-using, multi-step tasks. Every score is sourced and citable — see the provenance notes below before comparing across rows, harness configurations, or protocol versions.
| Model | ||
|---|---|---|
| Claude Opus 4.8 anthropic/claude-opus-4-8 | 88.5% 🏢 ↗ 2026-05
· multi-agent | 83.4% 🏢 ↗ 2026-05 |
| Claude Sonnet 5 anthropic/claude-sonnet-5 | 86.6% 🏢 ↗ 2026-06
· multi-agent | 81.2% 🏢 ↗ 2026-06 |
| Gemini 3.1 Pro google-deepmind/gemini-3.1-pro | 85.9% 🏢 ↗ 2026-02
· search + python + browse | — |
| GPT-5.5 openai/gpt-5.5 | 84.4% 🏢 ↗ 2026-04 | 78.7% 🏢 ↗ 2026-04 |
| MiniMax M3 minimax/m3 | 83.5% 🏢 ↗ 2026-06 | 70.1% 🏢 ↗ 2026-06 |
| Claude Sonnet 4.6 anthropic/claude-sonnet-4-6 | 76.2% 🏢 ↗ 2026-06 | 78.5% 🏢 ↗ 2026-06 |
| Claude Fable 5 anthropic/claude-fable-5 | — | 85.0% 🏢 ↗ 2026-06 |
| Gemini 3.5 Flash google-deepmind/gemini-3.5-flash | — | 78.4% 🏢 ↗ 2026-05 |
Sourcing and comparability
Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.
What is BrowseComp: BrowseComp (Wei et al., OpenAI, April 2025) tests an agent's ability to persistently browse the web to locate hard-to-find, entangled facts. Scores are harness-dependent — Anthropic, OpenAI, and Google DeepMind each disclose the tool configuration used (e.g. single-agent vs multi-agent), and only same-harness scores are directly comparable.
What is OSWorld-Verified: OSWorld-Verified (Xie et al., 2024) evaluates an agent's ability to complete real-world computer tasks — editing documents, browsing the web, managing files — by controlling a live Ubuntu virtual machine via mouse and keyboard. Protocol details (screen resolution, action-step cap, run count) materially affect scores; labs have retroactively corrected historical numbers after harness bug fixes, so only matching-protocol scores are comparable.
GAIA is legacy — not shown in this table
This table previously tracked GAIA (Mialon et al., 2023). It was replaced in 2026-07 as the anchor benchmark because no major lab self-reports a bare-model GAIA score for any current-gen model — a research pass found none, across the official leaderboard, every major vendor's own materials, and the paused Princeton HAL leaderboard. The registry keeps 2 historical GAIA entries (GPT-4 and GPT-4 Turbo, scored against GAIA's original 2023 paper) for the historical record, but since neither has a BrowseComp or OSWorld-Verified score, they aren't shown in this current-gen-focused table. See the modelglass-agentic repo for the full historical record.
Data is growing: This table is sourced from modelglass-agentic, a registry of agentic-task benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.