GDPval benchmarks

GDPval-AA v2.1 ranks models on real knowledge-work deliverables — documents, spreadsheets, slides, diagrams — by blind pairwise comparison rather than a fixed answer key. Every score on this leaderboard is sourced and citable — see the provenance notes below before comparing across rows or reasoning-effort configurations.

Frequently asked questions

What is GDPval-AA v2.1?
GDPval-AA v2.1 is Artificial Analysis's evaluation built on GDPval (OpenAI, 2025), a set of real-world knowledge-work tasks — a 220-task gold subset drawn from 1,320 tasks total — developed with industry professionals across finance, healthcare, legal, and other professional domains covering the top sectors contributing to U.S. GDP. Models work in an agentic loop with shell and web access and produce actual deliverables (documents, spreadsheets, slides, diagrams), which are compared blind, two at a time. The score is an Elo rating, not a pass-rate. Artificial Analysis anchors the v2.1 scale by pinning DeepSeek V4.1 Flash (max) at 1,600 and freezes it, so a number tells you how a model's deliverables compare with other models' on the same scale, not with a human professional's.
Which models are currently tracked?
19 models: Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1, Claude Opus 5, Grok 4.7, Grok 4.6, Claude Fable 5, GPT-5.6 Sol, Kimi K3, Claude Sonnet 5, GPT-5.6 Luna, DeepSeek V4-Pro, Claude Opus 4.8, GLM-5.2, GPT-5.5, MiniMax M3, Claude Sonnet 4.6, Gemini 3.5 Flash, MiMo-V2.5-Pro.
Which model scores highest on GDPval-AA v2.1?
Claude Opus 5.5, at 1,846 Elo as of 2026-09.
How often is the data updated?
This page was last rebuilt September 29, 2026 from the modelglass-gdpval registry. Scores are appended as new evaluations are verified — an existing result is never edited or deleted, only superseded by a newer entry.
Where does the score come from?
Every score on this page is an independent (🔬) evaluation administered by Artificial Analysis, not a lab's own self-reported figure. Some labs now quote Artificial Analysis's figures in their launch materials (Anthropic's Claude Opus 5.5 system card, 2026-09-22, cites GDPval-AA v2.1), but every score here is taken from Artificial Analysis directly. Each model is cited at its best-confirmed reasoning-effort configuration (disclosed under the score date); scores at different configurations aren't directly comparable. See the sourcing notes below for the current state of model coverage.

Who this is for

This page is built for business and operations leaders assessing which AI models handle real knowledge-work deliverables best — documents, spreadsheets, slides, and similar outputs — across roles like finance, healthcare, and legal, not abstract test questions. GDPval was designed by OpenAI explicitly as a practical roadmap for identifying which workflows can be augmented by AI; GDPval-AA v2.1 adds an independent, blind-comparison ranking of models on those tasks rather than a lab's own claim.

Model
Claude Opus 5.5
anthropic/claude-opus-5-5
1,846 Elo 🔬 ↗
2026-09
Claude Sonnet 5.5
anthropic/claude-sonnet-5-5
1,844 Elo 🔬 ↗
2026-09
Claude Fable 5.1
anthropic/claude-fable-5-1
1,735 Elo 🔬 ↗
2026-09
Claude Opus 5
anthropic/claude-opus-5
1,708 Elo 🔬 ↗
2026-09
Grok 4.7
xai/grok-4.7
1,695 Elo 🔬 ↗
2026-09
Grok 4.6
xai/grok-4.6
1,632 Elo 🔬 ↗
2026-09
Claude Fable 5
anthropic/claude-fable-5
1,595 Elo 🔬 ↗
2026-09
GPT-5.6 Sol
openai/gpt-5.6-sol
1,588 Elo 🔬 ↗
2026-09
Kimi K3
moonshot/kimi-k3
1,524 Elo 🔬 ↗
2026-09
Claude Sonnet 5
anthropic/claude-sonnet-5
1,449 Elo 🔬 ↗
2026-09
GPT-5.6 Luna
openai/gpt-5.6-luna
1,443 Elo 🔬 ↗
2026-09
DeepSeek V4-Pro
deepseek/deepseek-v4-pro
1,441 Elo 🔬 ↗
2026-09
Claude Opus 4.8
anthropic/claude-opus-4-8
1,438 Elo 🔬 ↗
2026-09
GLM-5.2
zhipu/glm-5.2
1,358 Elo 🔬 ↗
2026-09
GPT-5.5
openai/gpt-5.5
1,336 Elo 🔬 ↗
2026-09
MiniMax M3
minimax/m3
1,230 Elo 🔬 ↗
2026-09
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
1,220 Elo 🔬 ↗
2026-09
Gemini 3.5 Flash
google-deepmind/gemini-3.5-flash
1,185 Elo 🔬 ↗
2026-09
MiMo-V2.5-Pro
xiaomi/mimo-v2.5-pro
1,107 Elo 🔬 ↗
2026-09

See how benchmark data like this can power a cost-aware routing setup

See how →

Why we track this

We track GDPval-AA v2.1 because it's a genuine differentiator between models — not a vanity metric. Used alongside the other benchmarks on this page and pricing data elsewhere in the registry, it's designed to help you find the most appropriate model for your specific use case, not just the highest-scoring one.

Sourcing and comparability

Source types: 🏢 vendor-reported scores come directly from the model developer and may use proprietary evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs. 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications. Every score on this page is 🔬 independent — see below.

About the GDPval-AA v2.1 benchmark

GDPval-AA v2.1 builds on GDPval (OpenAI, 2025), a set of ~220 real-world knowledge-work tasks developed with industry professionals across finance, healthcare, legal, and other professional domains. Models work through each task in an agentic loop (Artificial Analysis's Stirrup harness, with shell and web access) and hand in actual deliverables — documents, spreadsheets, slides, diagrams. Those deliverables are compared blind, two at a time, and the results become an Elo rating, not a pass-rate.

How to read a v2.1 score

An Elo number only means something next to other numbers on the same scale. Artificial Analysis fixes the v2.1 scale by pinning one model, DeepSeek V4.1 Flash (max), at 1,600, and fits the ratings with a Crowd-BT model. The scale is frozen, so a model's rating shouldn't drift as new models join the pool. That's a change from v2, where published figures drifted as the pool grew. Read the gaps between models, not the absolute figure: a higher rating means judges preferred that model's deliverables to the other models' more consistently. There is no human-expert reference point on v2.1, so a score doesn't say whether a model beats a professional (see Artificial Analysis's methodology ↗).

Why the version changed: until mid-September 2026 this page used GDPval-AA v2, anchored to a human-expert baseline of 1,000. Artificial Analysis then replaced v2 with v2.1 at the same address. Its methodology says v2.1 "changes only how the Elo scale is fixed" (same tasks, harness and judging) and that rank ordering is "largely preserved", but every number moved. We switched on 2026-09-23 and re-checked every model against the live v2.1 leaderboard. The old v2 figures are kept in the source registry for history but aren't shown, because they sit on a different scale.

Why this is sourced differently from Coding/Science/Agentic: Every other vertical on this site anchors on scores labs self-report in their own system cards. GDPval-AA is run independently by Artificial Analysis instead. Some labs now quote its figures in their own launch materials (Anthropic's Claude Opus 5.5 system card, 2026-09-22, cites GDPval-AA v2.1), but every score here comes from Artificial Analysis directly. Every score here is cited as an independent evaluation, not a vendor claim, and each model is cited at its best-confirmed reasoning-effort configuration (disclosed under the score date) — scores at different configurations are not directly comparable.

What a GDPval task actually looks like

The tasks aren't quiz questions — they're the kind of assignment a professional gets handed on a Monday morning. GDPval's authors, industry practitioners averaging 14 years of experience, needed an average of about 7 hours per task to complete them themselves (median 5 hours; the longest run to multiple weeks), and each one ends in a real file: a spreadsheet, a slide deck, a legal memo, a diagram. Three examples from the openly published 220-task subset:

Accountant / auditor (Professional, Scientific & Technical Services). Given one reference workbook of raw tour income, costs and per-country tax data, build a profit-and-loss report for a seven-stop European concert tour: revenue itemised by city, foreign withholding applied at each country's rate (UK 20%, France 15%, Spain 24%, Germany 15.825%), everything converted to USD, expenses grouped into four categories, net income. Deliverable: a formatted Excel workbook, graded against a ~30-line checklist of specific figures (total net revenue must come to $852,428 USD) and formatting rules.

Audio & video technician (Information). From a written description of a five-piece touring band's monitoring and input/output needs, produce a one-page landscape PDF stage plot — icons for every amp, DI box, mic, monitor wedge and in-ear split positioned on stage, plus side-by-side numbered Input and Output lists, with the wedges numbered counterclockwise from stage right. Deliverable: a single-page PDF diagram. (Artificial Analysis publishes the actual PDFs different models turned in for this one.)

Residency program coordinator (Health Care & Social Assistance). From a spreadsheet of surgical case logs for an otolaryngology residency, establish per-year benchmarks — mean and standard deviation of each ACGME "key indicator" procedure — from the graduating cohort, flag any current resident more than two standard deviations below the mean on any measure, then draft the email to the program director summarising who was flagged and on which measures. Deliverable: an Excel workbook plus a Word email.

What the score means in practice: a higher GDPval-AA v2.1 rating means blind judges preferred that model's finished file over other models' on the same task more consistently. That's a statement about relative first-draft quality, not a promise of hands-off automation. Preference isn't accuracy either: judges picked the deliverable they preferred, not necessarily the one a specialist would find error-free. GDPval's own authors note that once you add the time a professional spends reviewing and fixing the output, the real-world speed-up is far smaller than the raw generation speed (roughly 1.0–1.2×, per arXiv 2510.04374). Reasoning-effort configuration matters too: the same model scores materially differently at medium, high or max effort, so each row is cited at its best-confirmed configuration.

19 models covered — more coverage to follow

Launched 2026-08-07 with four models and grown since. On 2026-09-23 every model was re-sourced from Artificial Analysis's GDPval-AA v2.1 leaderboard, Claude Opus 5.5 and Claude Fable 5.1 were added, and Grok 4.6 and 4.7 returned (their scores had been v2.1 figures in the old v2 column, so they were briefly removed). We only add a model once it also has a pricing entry, so several v2.1-rated models (GPT-6 Astra and Sol, GLM-5.3, Qwen3.8 Max, Muse Spark 1.3, MiMo-V2.6-Pro) are waiting on that. The leaderboard covers 250+ models. See the modelglass-gdpval repo for the current state of coverage.

Data is growing: This table is sourced from modelglass-gdpval, a registry of GDPval-AA knowledge-work benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.

Related pages