Last refreshed 2026-09-27 · one company's GTM stack · names changed

Two years of building with models. Then we started grading them.

Public charts score general skill. We needed to know which model handles our weekly work. So we built that test from our own extracts and graded it with code.

01 · TWO YEARS BUILDING

An AI-forward team with an open-model bet

Our team builds with models every day. We have done it for two years. Agents draft our reports, check our joins, and read our call notes.

Early on we made a deliberate bet on open models. We wanted low cost and full control. We wanted to switch models without rewriting our work.

That bet only pays if we can tell which model does the work well. Instinct could not tell us. A bench could.

02 · THE GAP

Nothing out there graded our work

Before this bench, our picks ran on instinct. We read release posts. We checked Vellum and OpenRouter charts. We asked around.

Those charts never graded RevOps work. None tested a rollforward. None tested a hygiene audit. None tested a forecast against messy extracts.

General scores did not predict who handled our data well. The only fix was a test built from our own jobs.

03 · WHAT WE BUILT

About 40 jobs from our own instance

We pulled about 40 jobs from our own instance. Forecasts, tiering, hygiene audits, renewal risk, the monthly close, call notes, lost-deal codes, joins across tools. Each hands the model real extracts and asks for the deliverable an operator produces each week.

We cap the set at 120 models. That covers 20 to 25 percent of active open models. We select by recency, quality, and performance. Frontier models stay in as the baseline.

The set changes as our work changes. New jobs replace old ones when the team stops doing them.

Task group

Forecasting

  • Model applies stage weights to open pipeline. Math must match to the dollar.
  • q3-weighted-forecast: commit and upside splits checked separately.
  • Stale close dates flagged, not silently included.
Task group

Pipeline quality

  • pipeline-tiering: each deal lands in the right tier. None skipped.
  • stage-hygiene-audit: stuck deals named with days in stage.
  • Blank owners counted, not dropped.
Task group

Renewals and risk

  • renewal-risk-conflicting-dates: both dates quoted when systems disagree.
  • Risk call follows the earlier date, not the later one.
  • No invented renewal owner. Whitelist enforced.
Task group

Revenue close

  • arr-rollforward-reconciliation: open plus adds minus churn equals close.
  • Every line item ties to the extract.
  • Off by a dollar on churn fails the check.
Task group

Call notes

  • call-transcript-extraction: facts match the transcript word for word.
  • Dates and amounts exact. Paraphrased numbers fail.
  • Claims stay with the right speaker.
Task group

Win and loss

  • closed-lost-classification: reason code matches the rubric.
  • Competitor named only if the notes name one.
  • Thin notes mean saying so, not guessing.
Task group

Data joins

  • gong-hubspot-join-integrity: dead deal ids kept out of the join.
  • Orphan calls listed apart, not forced in.
  • Row counts tie on both sides.
Task group

Ops honesty

  • Refusals count as correct when the data cannot answer.
  • Made-up names or accounts zero the response.
  • One rubric scored all 119 models. No exceptions.

Scope, and how to read the numbers

WhatDetail
RunFirst run 2026-09-05 · refreshed 2026-09-27 · same 40-task set and rubric
Models119 · 106 finished all 40 tasks · 13 partial, marked and ungraded on unrun tasks
Reference modelanthropic/claude-sonnet-5
GradingFixed Python against frozen answer files · no model judged another
DataOne company's GTM extracts · companies aliased · people renamed · zero grading drift
Compare ruleScores compare inside this run and rubric version only

Basis: 40 tasks, fixed rubric, one run. A fabrication flags the response. Partials are never scored on unrun tasks.

04 · PROOF IT WORKS

The bench picked Vera's stack

Vera is our RevOps agent. It runs on open models, and this report chose them. We implemented Jev, Typesafe's AI model, to deterministically route strategy and judgment work to meta/muse-spark-1.3. Routine mechanical work routes to z-ai/glm-5.3-flashx. The cheapest lookups go to tencent/hy3. Anything the router doubts falls back to Muse. We keep a short list of favorites and rotate as needed.

On the same 50 asks, routing offloaded 42% of turns from Muse and cut spend 30%. The mix came from these pages, and the meter confirms it.
The top pick, deepseek-pro-latest, scored 0.9846 for the full suite. Twenty-one models passed with top-tier efficiency. The priciest run, gpt-5.4-pro, ranks mid-table.

Twenty-two of 119 models invented an entity at least once. A name whitelist caught each case. Fabrication here is rare, and it is guardrail-able.

This page is not a demo. It is the sheet Vera reads before it spends our money.

05 · MONTHLY CADENCE

Work evolves, so picks evolve

The report is cumulative. We never re-run a model that is already on it. Each cycle runs new and stale models only. Past the 120 cap, we trim the lowest open scorers. Frontier baselines and the reference model stay.

CycleDetail
First run2026-09-05 to 2026-09-13 · 103 models · 4,066 graded answers
Refresh 2026-09-2421 models added (840 jobs, 818 graded) · 5 lowest open models trimmed · 0 re-run
Now119 models · 4,684 graded answers

The model registry and the per-cycle spend log are kept internally. The run-by-run comparison of tokens, cost, time, and failures stays off this site.

06 · BUILD YOUR OWN

Your data beats any public chart

The bench harness is open source. Runner, scorers, name whitelist, reviewer pages, one example suite. Swap in your extracts and your rules. Plan on about a week of work.

Get it here: github.com/amaniphipps/bonuslybench-starter. You supply the model credentials.

About me

Amani Phipps, Revenue Architect at Bonusly

I saw this problem in my work. Others shared it too. So I helped build this.

Happy to talk if you want to go deeper.

Connect on LinkedIn →