GDPval benchmarks

Knowledge-work capability on real deliverables — documents, spreadsheets, slides, diagrams — measured against a human-expert baseline rather than a fixed answer key. Every score is sourced and citable — see the provenance notes below before comparing across rows or reasoning-effort configurations.

Model
Claude Opus 5
anthropic/claude-opus-5
1,846 Elo (+846 vs. human) 🔬
2026-08
Claude Sonnet 5
anthropic/claude-sonnet-5
1,600 Elo (+600 vs. human) 🔬
2026-08
Claude Opus 4.8
anthropic/claude-opus-4-8
1,588 Elo (+588 vs. human) 🔬
2026-08
GPT-5.5
openai/gpt-5.5
1,490 Elo (+490 vs. human) 🔬
2026-08
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
1,375 Elo (+375 vs. human) 🔬
2026-08
Gemini 3.5 Flash
google-deepmind/gemini-3.5-flash
1,344 Elo (+344 vs. human) 🔬
2026-08

Sourcing and comparability

Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications. Every score on this page is 🔬 independent — see below.

What is GDPval-AA v2: GDPval-AA v2 builds on GDPval (OpenAI, 2025), a set of ~220 real-world knowledge-work tasks developed with industry professionals across finance, healthcare, legal, and other professional domains. Models produce actual deliverables — documents, spreadsheets, slides, diagrams — which are graded via blind pairwise comparison against a human-expert baseline, independently administered by Artificial Analysis. The score is an Elo rating, not a pass-rate: 1,000 Elo represents the human-expert baseline itself, so a model above 1,000 was preferred over a human professional's output more often than not in blind comparison.

Why this is sourced differently from Coding/Science/Agentic: Every other vertical on this site anchors on scores labs self-report in their own system cards. GDPval-AA v2 doesn't appear in any lab's own materials checked so far (confirmed against Claude Opus 5's launch system card, 2026-07-25) — it's run independently by Artificial Analysis instead. Every score here is cited as an independent evaluation, not a vendor claim, and each model is cited at its best-confirmed reasoning-effort configuration (disclosed under the score date) — scores at different configurations are not directly comparable.

6 models covered — more coverage to follow

Launched 2026-08-07 with four models; Claude Opus 5 and Claude Opus 4.8 were backfilled 2026-08-08. The full Artificial Analysis leaderboard covers 189 models — there's more room to grow. See the modelglass-gdpval repo for the current state of coverage.

Data is growing: This table is sourced from modelglass-gdpval, a registry of GDPval-AA v2 knowledge-work benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.