Leaderboard
Sort by any column. GLM-5.1 ties Opus 4.8 within 1.4pp at 12.5× lower cost — the real story is cost vs capability.
| Model | Pass rate↓ | Avg cost / run | Avg turns | Tool calls / run | Tool error rate | Total spend |
|---|---|---|---|---|---|---|
Claude Opus 4.8 anthropic/claude-opus-4.8 | 93.5% 201/215 | $0.0666 | 2.7 | — | — | $14.32 |
GPT-5.5 openai/gpt-5.5 | 93.0% 200/215 | $0.0278 | 2.5 | — | — | $5.97 |
GLM-5.1 z-ai/glm-5.1 | 92.1% 198/215 | $0.0053 | 2.9 | — | — | $1.15 |
Claude Sonnet 4.6 anthropic/claude-sonnet-4.6 | 91.6% 197/215 | $0.0324 | 2.6 | — | — | $6.97 |
Qwen3-235B qwen/qwen3-235b-a22b | 89.3% 192/215 | $0.0058 | 2.4 | — | — | $1.25 |
GPT-5 mini openai/gpt-5-mini | 88.4% 190/215 | $0.0036 | 2.5 | 3.0 | 33.3% | $0.78 |
Gemma 4 31B google/gemma-4-31b-it | 85.6% 184/215 | $0.0007 | 2.6 | 3.7 | 27.3% | $0.15 |
DeepSeek V4 Pro deepseek/deepseek-v4-pro | 85.1% 183/215 | $0.0094 | 3.2 | — | — | $2.02 |
Gemini 3.5 Flash google/gemini-3.5-flash | 81.4% 175/215 | $0.0162 | 3.8 | — | — | $3.48 |
Heatmap
Model × domain. Each cell aggregates all scenarios in that domain. Click to expand.
Cost vs pass rate
The publishable chart. Same vertical position = same accuracy; pick the leftmost. GLM-5.1 (92.1%) sits a whisker below Opus 4.8 (93.5%) — at 12.5× lower cost.
Cost vs pass rate
Each dot = one model. Log-scale cost axis. Models on the same horizontal line deliver the same accuracy — the leftmost are the value picks.
Domain breakdown
Where do models actually differentiate? Look at the short bars.
Pass rate by domain
Low bars = where models actually differentiate. High bars = scenarios too easy to discriminate.
What we're testing
Every scenario, grouped by domain. Click any row to see the exact prompt, system instructions, and what verify accepts as correct.
E-commerce
5 scenarios
Financial
8 scenarios
BI / Analytics
6 scenarios
Customer Service
8 scenarios
Marketing
8 scenarios
Adversarial (Safety)
10 scenarios