BridgePilot

Benchmark

9 models · 43 of 45 scenarios measured · 1,720 pass · 215 fail · 90 API errored (excluded from pass rate) · last aggregate 6/12/2026, 2:13:46 AM

Leaderboard

Sort by any column. GLM-5.1 ties Opus 4.8 within 1.4pp at 12.5× lower cost — the real story is cost vs capability.

ModelPass rateAvg cost / runAvg turnsTool calls / runTool error rateTotal spend
Claude Opus 4.8
anthropic/claude-opus-4.8
93.5%
201/215
$0.06662.7$14.32
GPT-5.5
openai/gpt-5.5
93.0%
200/215
$0.02782.5$5.97
GLM-5.1
z-ai/glm-5.1
92.1%
198/215
$0.00532.9$1.15
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
91.6%
197/215
$0.03242.6$6.97
Qwen3-235B
qwen/qwen3-235b-a22b
89.3%
192/215
$0.00582.4$1.25
GPT-5 mini
openai/gpt-5-mini
88.4%
190/215
$0.00362.53.033.3%$0.78
Gemma 4 31B
google/gemma-4-31b-it
85.6%
184/215
$0.00072.63.727.3%$0.15
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
85.1%
183/215
$0.00943.2$2.02
Gemini 3.5 Flash
google/gemini-3.5-flash
81.4%
175/215
$0.01623.8$3.48

Heatmap

Model × domain. Each cell aggregates all scenarios in that domain. Click to expand.

Pass
Mixed
Fail
Never ran
Click a cell to drill into scenarios
E-commerce
Financial
BI
Cust. Service
Marketing
Adversarial
Claude Opus 4.8
GPT-5.5
GLM-5.1
Claude Sonnet 4.6
Qwen3-235B
GPT-5 mini
Gemma 4 31B
DeepSeek V4 Pro
Gemini 3.5 Flash

Cost vs pass rate

The publishable chart. Same vertical position = same accuracy; pick the leftmost. GLM-5.1 (92.1%) sits a whisker below Opus 4.8 (93.5%) — at 12.5× lower cost.

Cost vs pass rate

Each dot = one model. Log-scale cost axis. Models on the same horizontal line deliver the same accuracy — the leftmost are the value picks.

vendors:
anthropic
openai
google
z-ai
deepseek
qwen

Domain breakdown

Where do models actually differentiate? Look at the short bars.

Pass rate by domain

Low bars = where models actually differentiate. High bars = scenarios too easy to discriminate.

What we're testing

Every scenario, grouped by domain. Click any row to see the exact prompt, system instructions, and what verify accepts as correct.

E-commerce

5 scenarios

100.0%
domain pass rate

Financial

8 scenarios

66.3%
domain pass rate

BI / Analytics

6 scenarios

88.1%
domain pass rate

Customer Service

8 scenarios

100.0%
domain pass rate

Marketing

8 scenarios

79.2%
domain pass rate

Adversarial (Safety)

10 scenarios

96.2%
domain pass rate