Last refreshed 2026-09-27 · 119 models · 40 tasks · graded by code, not by a model

Findings: cheap open models did this work as well as the expensive ones.

119 models ran 40 revenue-operations jobs drawn from one company's GTM extracts. Fixed code graded all 4,684 answers. Price did not predict accuracy.

01 · THE WRITEUP

The full story, with receipts

How a leaderboard habit became a 119-model bench, what the numbers say, the modality rule, the Vera live-fire experiments with screenshots, the stack that runs it, and which path fits your org. Twelve sections, every figure traceable to the data on this site.

02 · VERDICT

Price did not buy accuracy on this work

The top score belongs to deepseek/deepseek-pro-latest at 0.9846, with 37 of 40 perfect tasks, for $0.99 across the full suite. The priciest suite in the run, openai/gpt-5.4-pro at $131.87, scored 0.9421. That is roughly 133 times the spend for a lower score, and 16 of the 21 value kings named below outscore it.

Six models share the best mark of 36 of 40 perfect tasks. No model cleared all 40. The spread opened on tasks that mix arithmetic with judgment: which deals to exclude, what counts as stale, when to say the data cannot answer.

03 · VALUE SPLIT

Twenty-one models beat 0.93 at two dollars or less

Full-coverage models only, ranked by mean.

#ModelMean
1deepseek/deepseek-pro-latest0.9846
2z-ai/glm-5.3-flashx0.9781
3xiaomi/mimo-v2.5-pro0.9738
4qwen/qwen3.8-flash0.9692
5z-ai/glm-5.3-flash0.9679
6z-ai/glm-5.10.9679
7tencent/hy3-preview0.9677
8tencent/hy30.9671

Top 8 of the 21 value kings. Thirteen more full-coverage models score at least 0.93 at $2 or less.

The dollar spread is plain. The cheapest row costs $0.28 at a 0.9679 mean. The top row costs $0.99 at 0.9846. All eight sit under $2, while the priciest run in the report costs $131.87 and scores below each of them.

04 · SPEED AXIS

Fast and accurate exists, and so does slow and accurate

Median response time ran from under 20 seconds to over 40 minutes per task. Speed reads as its own axis, separate from accuracy and from price.

tencent/hy3-preview averaged 82 seconds a task at a 0.9677 mean and $0.59 for the suite. qwen/qwen3.8-flash averaged 111 seconds a task at 0.9692 and $0.34. Both are cheap, fast, and near the top of the board.

The contrast is tencent/hy3: about 27 minutes a task on average, yet 0.9671 at $0.33. Same score band, same low price, very different wait. For embedded workflows, that wait is the buying decision.

05 · THIS CYCLE

What moved in the refresh

deepseek/deepseek-pro-latest is the new top at 0.9846, ahead of the prior best of 0.9773 from meta/muse-spark-1.1. The refresh added 21 models and trimmed 5. The trimmed five sat at the bottom:

The 21 new models averaged $4.19 each against $7.50 in the initial run, with mean quality up at 0.940. The report now covers 119 models and 4,684 graded answers at $860.73 total. Nothing already on the report was re-run.

About me

Amani Phipps, Revenue Architect at Bonusly

I saw this problem in my work. Others shared it too. So I helped build this.

Happy to talk if you want to go deeper.

Connect on LinkedIn →