GDPval benchmarks
Knowledge-work capability on real deliverables — documents, spreadsheets, slides, diagrams — measured against a human-expert baseline rather than a fixed answer key. Every score is sourced and citable — see the provenance notes below before comparing across rows or reasoning-effort configurations.
| Model | |
|---|---|
| Claude Opus 5 anthropic/claude-opus-5 | 1,846 Elo (+846 vs. human) 🔬 ↗ 2026-08 |
| Claude Sonnet 5 anthropic/claude-sonnet-5 | 1,600 Elo (+600 vs. human) 🔬 ↗ 2026-08 |
| Claude Opus 4.8 anthropic/claude-opus-4-8 | 1,588 Elo (+588 vs. human) 🔬 ↗ 2026-08 |
| GPT-5.5 openai/gpt-5.5 | 1,490 Elo (+490 vs. human) 🔬 ↗ 2026-08 |
| Claude Sonnet 4.6 anthropic/claude-sonnet-4-6 | 1,375 Elo (+375 vs. human) 🔬 ↗ 2026-08 |
| Gemini 3.5 Flash google-deepmind/gemini-3.5-flash | 1,344 Elo (+344 vs. human) 🔬 ↗ 2026-08 |
Sourcing and comparability
Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications. Every score on this page is 🔬 independent — see below.
What is GDPval-AA v2: GDPval-AA v2 builds on GDPval (OpenAI, 2025), a set of ~220 real-world knowledge-work tasks developed with industry professionals across finance, healthcare, legal, and other professional domains. Models produce actual deliverables — documents, spreadsheets, slides, diagrams — which are graded via blind pairwise comparison against a human-expert baseline, independently administered by Artificial Analysis. The score is an Elo rating, not a pass-rate: 1,000 Elo represents the human-expert baseline itself, so a model above 1,000 was preferred over a human professional's output more often than not in blind comparison.
Why this is sourced differently from Coding/Science/Agentic: Every other vertical on this site anchors on scores labs self-report in their own system cards. GDPval-AA v2 doesn't appear in any lab's own materials checked so far (confirmed against Claude Opus 5's launch system card, 2026-07-25) — it's run independently by Artificial Analysis instead. Every score here is cited as an independent evaluation, not a vendor claim, and each model is cited at its best-confirmed reasoning-effort configuration (disclosed under the score date) — scores at different configurations are not directly comparable.
6 models covered — more coverage to follow
Launched 2026-08-07 with four models; Claude Opus 5 and Claude Opus 4.8 were backfilled 2026-08-08. The full Artificial Analysis leaderboard covers 189 models — there's more room to grow. See the modelglass-gdpval repo for the current state of coverage.
Data is growing: This table is sourced from modelglass-gdpval, a registry of GDPval-AA v2 knowledge-work benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.