Open models did this work as well as the expensive ones. That held up across 119 models, 40 real tasks, and 4,684 graded answers. The most useful findings are not about models at all.
My old process was a models page and a benchmarking page. See what is popular and new, then go try it. That is how a Tencent model named Hy3 ended up doing real work at our company. It showed up near the top of everything, I used it, and it was genuinely good.
What the leaderboard never gave me was a so-what. I kept using it, watched the actual cost, noticed it ran more verbose than what I already had, and started running small side-by-side comparisons against the four or five models actually in rotation: GLM, Qwen, DeepSeek, and the Anthropic and OpenAI models I was splitting between.
I could see something real in the open field. I could also see several hundred more models I had not touched, and testing them one at a time was not a plan. That gap is BonuslyBench.
The first design question was how to make the results trustworthy. The answer was to test against our own skills rather than benchmark-style tasks. By that point we had built well over a hundred skills for our AI instance, with a deliberate rollout behind them: foundation first, alignment on governance, a real jobs-to-be-done check before anything shipped.
It was not tidy from day one. We had a token-maxing month early on where we just explored, because people needed to be using this stuff before anyone could see where value was hiding. Once that phase did its job, the work slowed down and got specific: skills built for one team, solving a problem that team had already said was worth solving.
The bench itself is 40 revenue-operations tasks drawn from real GTM extracts: pipeline hygiene, forecasting, renewal risk, attribution, customer health. Five tests in each of eight categories. Every task ships with its source data and a plain-English question, plus ground truth computed by code from that same data. Answers are graded by fixed code, never by a model. A name whitelist catches invented accounts, owners, and numbers. That is what makes it possible to ask an honest question: is this model cheaper than what I already pay for the same job?
The headline result: open models did this work as well as the expensive ones. The top score belongs to deepseek/deepseek-pro-latest at 0.9846, with 37 of 40 tasks perfect, for $0.99 across the full suite. The priciest suite in the run, openai/gpt-5.4-pro at $131.87, scored 0.9421 with 32 of 40 perfect. That is roughly 133 times the spend for a lower score. Thirty-nine full-coverage models outscore it, and 16 of the 21 value kings named below outscore it.
Six models share the next-best mark of 36 perfect tasks. None cleared all 40. The spread opens on tasks that mix arithmetic with judgment: which deals to exclude, what counts as stale, when to say the data cannot answer the question.
| # | Model | Mean | Perfect | |
|---|---|---|---|---|
| 1 | deepseek/deepseek-pro-latest | 0.9846 | 37/40 | $0.99 |
| 2 | fireworks/ember-1 | 0.9796 | 36/40 | $7.31 |
| 3 | z-ai/glm-5.3-flashx | 0.9781 | 36/40 | $1.34 |
| 4 | meta/muse-spark-1.1 | 0.9773 | 35/40 | $2.17 |
| 5 | moonshotai/kimi-k3 | 0.9762 | 36/40 | $7.67 |
| 6 | xiaomi/mimo-v2.5-pro | 0.9738 | 36/40 | $0.69 |
| 7 | qwen/qwen3.8-flash | 0.9692 | 35/40 | $0.34 |
| 8 | z-ai/glm-5.3-flash | 0.9679 | 35/40 | $0.28 |
Speed turned out to be its own axis, separate from accuracy and price. tencent/hy3-preview averaged 82 seconds a task at 0.9677. tencent/hy3 took about 27 minutes a task at essentially the same score (0.9671). If someone is waiting on an answer in Slack, that difference decides the purchase.
Failure modes were specific. 22 of 119 models invented something: a name, an account, a renewal owner that was not in the data. The whitelist caught each case and the answer scored zero. That is the failure to guardrail first. Thirteen models ran short of the full 40, mostly slow responders timing out; those tasks sit ungraded and marked partial. We never score work a model did not do.
The top model on my list could not read a screenshot. DeepSeek is a text-only model. It leads the board on text work and it cannot process files, images, or video. I learned that one live: I sent it a screenshot of a bug the way I would send one to any coworker, and it just could not take it. Best model on the list, wrong model for that ten seconds of work.
This matters for anything that lives in Slack. People talk to an agent there the same way they talk to the person next to them, and the natural response to "here's what I'm seeing" is a screenshot. A model that cannot see it is out of the running for that job, no matter where it ranks.
So the model you use depends on your use case. And the line moves. Harnesses and IDEs can encode formats the model itself cannot take. A screenshot becomes a text extract, a PDF becomes pasted text, and a text-only model is back in the game. Hermes does this kind of preprocessing before the model ever sees the input, which softens the boundary the bench scores draw.
But keep one thing in mind. This bench is a benchmark on my work. My tasks, my data, my jobs-to-be-done. The scores tell you where I would start looking. For your own work, you have to try and see for yourself. No benchmark, mine included, picks the model for you.
The bench scores answers on paper. The other half of the testing happened live. Between refreshes we ran a series of experiments with Vera, our RevOps teammate. Slack is where Vera meets the company, but that is just the front door. Behind it sits the full stack: the skills, the governance, the documentation, and a weekly loop that reads RevenueBench and picks the model for each piece of work. For the experiments, we swapped a candidate model into that stack, handed it real requests from real people, and saw what came back. Same question, two models. Then we compared notes on the things that matter: did the answer hold up, how long did it take, and would a coworker actually want to read it.
The bench catches hallucinations and math mistakes. Vera catches everything a score cannot: the model that is right but takes four minutes in a channel where thirty seconds is the norm, or the one that writes a technically perfect summary nobody wants to read. A few high scorers did not survive real Slack traffic. A couple of mid-board models earned a permanent spot because of how they behaved in the product.
Both halves feed one decision. The bench narrows the field. Vera makes the final call. An agent now reads all of RevenueBench on a weekly loop and answers "what model suits this piece of work," and that is how dozens of models nobody would have tried on their own made it into rotation. This run's top scorer came out of that loop, not from a leaderboard page.
Three models at the top tell three different stories, so here is the full breakdown.
| Model | Full bench | 20-test subset | Avg latency | |
|---|---|---|---|---|
| fireworks/ember-1 | 0.9796 · 36/40 | 0.946 | $7.31 | 124.8s |
| meta/muse-spark-1.1 | 0.9773 · 35/40 | 0.949 | $2.17 | 40.2s |
| moonshotai/kimi-k3 | 0.9762 · 36/40 | 0.980 | $7.67 | 102.1s |
Kimi is my daily driver. It costs more than Muse ($7.67 against $2.17 for the full suite) and I pay it gladly. It topped the 20-test subset at 0.980, went clean on the mockup, and holds 36 perfect tasks on the full bench. Against the frontier models it replaces, the gap is not close. GPT-5.5-pro costs $123.62 for a lower score, more than sixteen times what Kimi costs. That is the whole trade in one line: slightly more than the cheapest good answer, a fraction of the famous ones.
Muse-Spark-1.1 is the speed pick. Forty seconds a task, fastest in the top tier, at $2.17. When someone is waiting in a channel, that is the call.
Ember-1 is the accuracy runner-up at 0.9796 with 36 perfect. It trailed the subset at 0.946, which is exactly why the subset exists. One number never tells you enough.
Neither number alone picks the model. The bench is the filter, the product is the verdict, and you need both or you will promote a model that is right on paper and wrong in the room.
For the builders who want the schematic, here is the whole pipeline. The harness is open source (bonuslybench-starter on GitHub, MIT): a runner, deterministic scorers, mock tests, and OpenRouter plus Hermes providers. Nothing in the grading path is a model. That is a design constraint, not a preference: "20 of 103 models fabricated at least one answer" is a defensible headline only because the grader is code a human can audit.
run_bench.py fans tasks out across models. A 21-model append run is about 480 jobs and three hours.score_all.py grades every answer with fixed checks. One final retry for stragglers, then timeouts are discarded and marked partial. Never scored as zero, never silently dropped.refresh_all.py regenerates every artifact from on-disk state: matrix, cost-value tiers, archetypes, recommendations, reviewer pages, dashboard charts, the research docx. The build script is the only writer. Twice, a hand-copied file shipped 404s across a live results surface; the bug class dies only when exactly one code path produces the artifact.verify_site_numbers.py fails loudly on drift between scored state and the hand-maintained story pages. It runs before every deploy. No green check, no ship.Vera itself runs Railway-only: the teammate, the weekly bench-read loop, and the model routing all live there. Local runs are cron-only. That split is deliberate. The product and the measurement stay on the same rails.
I am a heavier AI spender than almost anyone else at the company, which is not an accident: I build the tools most people here use. Building on open models has brought that cost down by a lot, not a little.
SignalForge, the first tool we built end to end, took a month and a half and roughly five to six times what Vera cost, mostly in testing and iteration. Vera took about two weeks. The difference was not the model. By the time I built Vera, the skills, the governance, and the documentation were already sitting there, ready to transfer over. Rebuild SignalForge today knowing everything we know now and it would still run three to four times what Vera cost.
That is the real return on infrastructure work: not the first tool it powers, but the second and third ones you build faster because of it. Skills, plugins, governance, documentation. That is what keeps a team able to move between tools and models without starting from zero.
Chasing the perfect open model and building agents around it before you understand your use case is how teams burn months and end up worse off than if they had stayed on Claude. We got conviction on this early: good data, good systems, documentation that means something, consistency across everything we built. That took months, and it is not something you build once and walk away from. Workflows get re-checked on a schedule, because the foundation slips if nobody keeps an eye on it.
If this research proved one thing more than any other, it is that your data matters more than which model you pick.
I am building Vera to work for every employee at the company, a large volume of calls across a lot of different jobs, all day, most days. At that volume, open models are genuinely more viable than at a fraction of the usage. The math changes with scale, in both directions.
Org size works differently. Mid-market and enterprise teams with several layers of decision-making usually need predictability more than the cheapest token. Open models fit a subset of orgs in that category, not the default. If you go anyway, a platform that manages a small, consistent set of models beats swapping between dozens yourself. The infrastructure is the part orgs underestimate.
One honest caveat: open-model costs do not run as flat month to month as a single-vendor bill. Usage patterns shift as you move between models in ways they do not when you are anchored to one. If predictability matters more to you than squeezing out the last dollar, Claude is a perfectly fine answer.
Findings are only useful if they tell someone what to do Monday morning. Here is my read, by org shape.
When: your infra is easy to change, your data is in good shape, and someone owns the stack. You can swap models, watch the numbers, and re-check on a schedule.
Payoff: this is where the full win lives. Best scores at the lowest cost, and honestly it is fun. Picking the model per job never gets old.
When: you want open-model pricing without running the bench yourself. A platform like Hoonify gives you a small, steady set of great models. Not all of them, just a few. Someone else handles the routing and the updates.
Payoff: most of the benefit, none of the tinkering. Fewer models to choose from, but that is the point.
When: your org needs a flat bill and a single vendor to hold accountable. Several layers of approvals, hundreds or thousands of employees, no appetite for costs that move month to month.
Payoff: Claude is a perfectly fine answer. Open-model costs shift as you move between models in ways a single-vendor bill does not. If that variance scares you more than the price tag, stay put.
Whichever path you take, three rules travel with you. Match the modality to the job before you compare scores. Keep a live-fire loop, because the bench narrows the field and the product makes the call. And guardrail first: 22 of 119 models invented something, and the whitelist is the only reason those answers scored zero instead of shipping.
Scores pick the shortlist. Living with the models picks the winners. These are mine, in the order I reach for them.
Kimi K3. The game changer. Our team had been running on Qwen and GLM models, which were fine, and then K3 showed up and it has just been incredible. It is my daily driver: 0.9762 on the full bench with 36 perfect tasks, top of the 20-test subset at 0.980, clean on the mockup. The full breakdown sits in section 5.
Muse Spark 1.1. This one shocked me. I did not expect it to score near the top, and then it also turned out to do incredible work at a lower cost than K3. $2.17 against $7.67 for the full suite, and the fastest in the top tier at 40 seconds a task. When someone is waiting in a channel, this is the call.
Hy3 and Hy4. The Tencent models hold a special place. They are not the best at all work. They are night and day cheaper and still do really good work. Hy3 is the one I keep coming back to: 36 perfect tasks at $0.33 for the suite, which outscores Haiku-4.5 (29 perfect at $2.34). I will likely replace Haiku with Hy3 where shavings make a pile, across the numerous CRM and system jobs we run.
Hoonify's GLM and Qwen. We ran their hosted GLM and Qwen models against the bench and they were solid. This is the managed path from section 10 working in practice: a few good models, none of the tinkering.
Pareto and Ember. Both on my list. Ember sits second overall at 0.9796, Pareto matches Kimi's 36 perfect at 0.9733. I picture ending up with a suite of five or so core open models per department, with the same three or four underneath everything I do.
DeepSeek. So good it is silly. Top of the board at 0.9846 for $0.99. Yes, I was bummed when I learned it could not see the files I shared in real work (section 4). Even so: fast, affordable, high quality. It earns the top spot.
The most valuable output is not the savings. It is the pulse check. Without this I would be paying whatever Claude charges because Claude is what we have always used, and never asking the question. Getting equal or better results out of open models keeps reminding me how much is actually moving right now, and that is worth watching whether or not it saves a dollar this month.
We will keep testing as long as there is something worth testing. This was the first full pass.
Live report: bonuslybench.com · Open harness: github.com/amaniphipps/bonuslybench-starter · All figures from summary_data.json, September 2026 run (119 models, 4,684 graded answers).