119 models ran 40 revenue-operations jobs drawn from one company's GTM extracts. Fixed code graded all 4,684 answers. Price did not predict accuracy.
How a leaderboard habit became a 119-model bench, what the numbers say, the modality rule, the Vera live-fire experiments with screenshots, the stack that runs it, and which path fits your org. Twelve sections, every figure traceable to the data on this site.
The top score belongs to deepseek/deepseek-pro-latest at 0.9846, with 37 of 40 perfect tasks, for $0.99 across the full suite. The priciest suite in the run, openai/gpt-5.4-pro at $131.87, scored 0.9421. That is roughly 133 times the spend for a lower score, and 16 of the 21 value kings named below outscore it.
Six models share the best mark of 36 of 40 perfect tasks. No model cleared all 40. The spread opened on tasks that mix arithmetic with judgment: which deals to exclude, what counts as stale, when to say the data cannot answer.
Full-coverage models only, ranked by mean.
| # | Model | Mean |
|---|---|---|
| 1 | deepseek/deepseek-pro-latest | 0.9846 |
| 2 | z-ai/glm-5.3-flashx | 0.9781 |
| 3 | xiaomi/mimo-v2.5-pro | 0.9738 |
| 4 | qwen/qwen3.8-flash | 0.9692 |
| 5 | z-ai/glm-5.3-flash | 0.9679 |
| 6 | z-ai/glm-5.1 | 0.9679 |
| 7 | tencent/hy3-preview | 0.9677 |
| 8 | tencent/hy3 | 0.9671 |
Top 8 of the 21 value kings. Thirteen more full-coverage models score at least 0.93 at $2 or less.
The dollar spread is plain. The cheapest row costs $0.28 at a 0.9679 mean. The top row costs $0.99 at 0.9846. All eight sit under $2, while the priciest run in the report costs $131.87 and scores below each of them.
Median response time ran from under 20 seconds to over 40 minutes per task. Speed reads as its own axis, separate from accuracy and from price.
tencent/hy3-preview averaged 82 seconds a task at a 0.9677 mean and $0.59 for the suite. qwen/qwen3.8-flash averaged 111 seconds a task at 0.9692 and $0.34. Both are cheap, fast, and near the top of the board.
The contrast is tencent/hy3: about 27 minutes a task on average, yet 0.9671 at $0.33. Same score band, same low price, very different wait. For embedded workflows, that wait is the buying decision.
deepseek/deepseek-pro-latest is the new top at 0.9846, ahead of the prior best of 0.9773 from meta/muse-spark-1.1. The refresh added 21 models and trimmed 5. The trimmed five sat at the bottom:
The 21 new models averaged $4.19 each against $7.50 in the initial run, with mean quality up at 0.940. The report now covers 119 models and 4,684 graded answers at $860.73 total. Nothing already on the report was re-run.
I saw this problem in my work. Others shared it too. So I helped build this.
Happy to talk if you want to go deeper.