Public charts score general skill. We needed to know which model handles our weekly work. So we built that test from our own extracts and graded it with code.
Our team builds with models every day. We have done it for two years. Agents draft our reports, check our joins, and read our call notes.
Early on we made a deliberate bet on open models. We wanted low cost and full control. We wanted to switch models without rewriting our work.
That bet only pays if we can tell which model does the work well. Instinct could not tell us. A bench could.
Before this bench, our picks ran on instinct. We read release posts. We checked Vellum and OpenRouter charts. We asked around.
Those charts never graded RevOps work. None tested a rollforward. None tested a hygiene audit. None tested a forecast against messy extracts.
General scores did not predict who handled our data well. The only fix was a test built from our own jobs.
We pulled about 40 jobs from our own instance. Forecasts, tiering, hygiene audits, renewal risk, the monthly close, call notes, lost-deal codes, joins across tools. Each hands the model real extracts and asks for the deliverable an operator produces each week.
We cap the set at 120 models. That covers 20 to 25 percent of active open models. We select by recency, quality, and performance. Frontier models stay in as the baseline.
The set changes as our work changes. New jobs replace old ones when the team stops doing them.
q3-weighted-forecast: commit and upside splits checked separately.pipeline-tiering: each deal lands in the right tier. None skipped.stage-hygiene-audit: stuck deals named with days in stage.renewal-risk-conflicting-dates: both dates quoted when systems disagree.arr-rollforward-reconciliation: open plus adds minus churn equals close.call-transcript-extraction: facts match the transcript word for word.closed-lost-classification: reason code matches the rubric.gong-hubspot-join-integrity: dead deal ids kept out of the join.| What | Detail |
|---|---|
| Run | First run 2026-09-05 · refreshed 2026-09-27 · same 40-task set and rubric |
| Models | 119 · 106 finished all 40 tasks · 13 partial, marked and ungraded on unrun tasks |
| Reference model | anthropic/claude-sonnet-5 |
| Grading | Fixed Python against frozen answer files · no model judged another |
| Data | One company's GTM extracts · companies aliased · people renamed · zero grading drift |
| Compare rule | Scores compare inside this run and rubric version only |
Basis: 40 tasks, fixed rubric, one run. A fabrication flags the response. Partials are never scored on unrun tasks.
Vera is our RevOps agent. It runs on open models, and this report chose them. We implemented Jev, Typesafe's AI model, to deterministically route strategy and judgment work to meta/muse-spark-1.3. Routine mechanical work routes to z-ai/glm-5.3-flashx. The cheapest lookups go to tencent/hy3. Anything the router doubts falls back to Muse. We keep a short list of favorites and rotate as needed.
Twenty-two of 119 models invented an entity at least once. A name whitelist caught each case. Fabrication here is rare, and it is guardrail-able.
This page is not a demo. It is the sheet Vera reads before it spends our money.
The report is cumulative. We never re-run a model that is already on it. Each cycle runs new and stale models only. Past the 120 cap, we trim the lowest open scorers. Frontier baselines and the reference model stay.
| Cycle | Detail |
|---|---|
| First run | 2026-09-05 to 2026-09-13 · 103 models · 4,066 graded answers |
| Refresh 2026-09-24 | 21 models added (840 jobs, 818 graded) · 5 lowest open models trimmed · 0 re-run |
| Now | 119 models · 4,684 graded answers |
The model registry and the per-cycle spend log are kept internally. The run-by-run comparison of tokens, cost, time, and failures stays off this site.
The bench harness is open source. Runner, scorers, name whitelist, reviewer pages, one example suite. Swap in your extracts and your rules. Plan on about a week of work.
Get it here: github.com/amaniphipps/bonuslybench-starter. You supply the model credentials.
I saw this problem in my work. Others shared it too. So I helped build this.
Happy to talk if you want to go deeper.