Agentic benchmarks

Quantitative agentic performance on real-world, tool-using, multi-step tasks. Every score is sourced and citable — see the provenance notes below before comparing across rows, harness configurations, or protocol versions.

Frequently asked questions

What is BrowseComp?
BrowseComp (Wei et al., OpenAI, April 2025) is a 1,266-question benchmark that tests an agent's ability to persistently browse the web to locate hard-to-find, entangled facts — the answers are short and easily verified, but finding them requires real, sustained search rather than a single lookup. Scores are harness-dependent: Anthropic, OpenAI, and Google DeepMind each disclose the tool configuration used (single-agent vs. multi-agent), and only same-harness scores are directly comparable. This page also tracks OSWorld-Verified, real computer-use tasks completed via a live Ubuntu virtual machine — see the column header and sourcing notes below.
Which models are currently tracked?
8 models: Claude Opus 4.8, Claude Sonnet 5, Gemini 3.1 Pro, GPT-5.5, MiniMax M3, Claude Sonnet 4.6, Claude Fable 5, Gemini 3.5 Flash.
Which model scores highest on BrowseComp?
Claude Opus 4.8, at 88.5% as of 2026-05.
How often is the data updated?
This page was last rebuilt July 9, 2026 from the modelglass-agentic registry. Scores are appended as new evaluations are verified — an existing result is never edited or deleted, only superseded by a newer entry.
Where does the score come from?
Every score below links to its source (the ↗ icon next to each result) and shows an icon for how it was obtained: 🏢 vendor-reported (the model developer's own announcement), 📊 public leaderboard (a public, independently-maintained evaluation run), 🔬 independent evaluation (a credible third party ran the benchmark), or 📄 published paper (a peer-reviewed or arXiv result). See the sourcing notes below the table for BrowseComp/OSWorld-Verified harness and protocol caveats, and the note on GAIA's retirement as this page's former anchor benchmark.

Who this is for

This page is built for teams building or evaluating autonomous agents that need to act, not just answer — persistently searching the web for hard-to-find facts, or operating a real desktop environment via mouse and keyboard. BrowseComp's own creators describe it as analogous to programming competitions for coding agents — an imperfect but useful stress test for a specific capability; OSWorld-Verified plays the same role for computer-use agents completing real, multi-step desktop tasks.

Model
Claude Opus 4.8
anthropic/claude-opus-4-8
88.5% 🏢
2026-05 · multi-agent
83.4% 🏢
2026-05
Claude Sonnet 5
anthropic/claude-sonnet-5
86.6% 🏢
2026-06 · multi-agent
81.2% 🏢
2026-06
Gemini 3.1 Pro
google-deepmind/gemini-3.1-pro
85.9% 🏢
2026-02 · search + python + browse
GPT-5.5
openai/gpt-5.5
84.4% 🏢
2026-04
78.7% 🏢
2026-04
MiniMax M3
minimax/m3
83.5% 🏢
2026-06
70.1% 🏢
2026-06
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
76.2% 🏢
2026-06
78.5% 🏢
2026-06
Claude Fable 5
anthropic/claude-fable-5
85.0% 🏢
2026-06
Gemini 3.5 Flash
google-deepmind/gemini-3.5-flash
78.4% 🏢
2026-05

See how benchmark data like this can power a cost-aware routing setup

See how →

Why we track this

We track BrowseComp because it's a genuine differentiator between models — not a vanity metric. Used alongside the other benchmarks on this page and pricing data elsewhere in the registry, it's designed to help you find the most appropriate model for your specific use case, not just the highest-scoring one.

Sourcing and comparability

Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.

What is BrowseComp: BrowseComp (Wei et al., OpenAI, April 2025) tests an agent's ability to persistently browse the web to locate hard-to-find, entangled facts. Scores are harness-dependent — Anthropic, OpenAI, and Google DeepMind each disclose the tool configuration used (e.g. single-agent vs multi-agent), and only same-harness scores are directly comparable.

What is OSWorld-Verified: OSWorld-Verified (Xie et al., 2024) evaluates an agent's ability to complete real-world computer tasks — editing documents, browsing the web, managing files — by controlling a live Ubuntu virtual machine via mouse and keyboard. Protocol details (screen resolution, action-step cap, run count) materially affect scores; labs have retroactively corrected historical numbers after harness bug fixes, so only matching-protocol scores are comparable.

GAIA is legacy — not shown in this table

This table previously tracked GAIA (Mialon et al., 2023). It was replaced in 2026-07 as the anchor benchmark because no major lab self-reports a bare-model GAIA score for any current-gen model — a research pass found none, across the official leaderboard, every major vendor's own materials, and the paused Princeton HAL leaderboard. The registry keeps 2 historical GAIA entries (GPT-4 and GPT-4 Turbo, scored against GAIA's original 2023 paper) for the historical record, but since neither has a BrowseComp or OSWorld-Verified score, they aren't shown in this current-gen-focused table. See the modelglass-agentic repo for the full historical record.

Data is growing: This table is sourced from modelglass-agentic, a registry of agentic-task benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.

Related pages