Agentic benchmarks

Quantitative agentic performance on real-world, tool-using, multi-step tasks. Every score is sourced and citable — see the provenance notes below before comparing across rows, harness configurations, or protocol versions.

Model
Claude Opus 4.8
anthropic/claude-opus-4-8
88.5% 🏢
2026-05 · multi-agent
83.4% 🏢
2026-05
Claude Sonnet 5
anthropic/claude-sonnet-5
86.6% 🏢
2026-06 · multi-agent
81.2% 🏢
2026-06
Gemini 3.1 Pro
google-deepmind/gemini-3.1-pro
85.9% 🏢
2026-02 · search + python + browse
GPT-5.5
openai/gpt-5.5
84.4% 🏢
2026-04
78.7% 🏢
2026-04
MiniMax M3
minimax/m3
83.5% 🏢
2026-06
70.1% 🏢
2026-06
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
76.2% 🏢
2026-06
78.5% 🏢
2026-06
Claude Fable 5
anthropic/claude-fable-5
85.0% 🏢
2026-06
Gemini 3.5 Flash
google-deepmind/gemini-3.5-flash
78.4% 🏢
2026-05

Sourcing and comparability

Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.

What is BrowseComp: BrowseComp (Wei et al., OpenAI, April 2025) tests an agent's ability to persistently browse the web to locate hard-to-find, entangled facts. Scores are harness-dependent — Anthropic, OpenAI, and Google DeepMind each disclose the tool configuration used (e.g. single-agent vs multi-agent), and only same-harness scores are directly comparable.

What is OSWorld-Verified: OSWorld-Verified (Xie et al., 2024) evaluates an agent's ability to complete real-world computer tasks — editing documents, browsing the web, managing files — by controlling a live Ubuntu virtual machine via mouse and keyboard. Protocol details (screen resolution, action-step cap, run count) materially affect scores; labs have retroactively corrected historical numbers after harness bug fixes, so only matching-protocol scores are comparable.

GAIA is legacy — not shown in this table

This table previously tracked GAIA (Mialon et al., 2023). It was replaced in 2026-07 as the anchor benchmark because no major lab self-reports a bare-model GAIA score for any current-gen model — a research pass found none, across the official leaderboard, every major vendor's own materials, and the paused Princeton HAL leaderboard. The registry keeps 2 historical GAIA entries (GPT-4 and GPT-4 Turbo, scored against GAIA's original 2023 paper) for the historical record, but since neither has a BrowseComp or OSWorld-Verified score, they aren't shown in this current-gen-focused table. See the modelglass-agentic repo for the full historical record.

Data is growing: This table is sourced from modelglass-agentic, a registry of agentic-task benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.