RevenueBench · Model Recommendations · September 2026
Which Models Fit Which Work
Work-type recommendations drawn from measured performance on real GTM tasks, deterministically scored. Match the model to the type of work you're doing — the data shows clear specialists.
📖 How to use this: find the section that matches your situation below. Start with the #1 model in that section, then test it on 20 to 40 of your own jobs before you commit. These scores come from one company's GTM data. They are a strong starting point, not a promise. Every model name links to its full answer-by-answer record.
Coming from Anthropic or OpenAI? Start Here
New to open models? These five match big-lab output on operational work at a fraction of the cost. Ranked by overall score, with a floor: the weakest category must still be strong.
Strong all-around open model — top-ranked on combined mean and cross-category breadth.
What the Job Arrives As Matters First
A model that cannot read images scores zero on screenshot work no matter how it ranks. Check what your jobs arrive as, then pick the strongest model that can take it. Each group below shows the top three full-coverage models by overall mean.
Media generalists (24 full-coverage models read this)
Jobs that arrive as video, audio, screenshots, or files. Slack threads with recordings, call videos, pasted clips.
Heavy multi-table reconciliation (rollforwards, weighted forecasts, cross-system audits): the best reporting models clearly separate from the pack on these tests. Do not hand scheduled pipeline math to a chat-tuned model by default.
Executive communication (digests, summaries, RFP answers): 35 models swept the communication tests. Same mark, very different bill.
Speed-sensitive embeds (CRM hooks, in-app enrichment): Qwen 3.8 Flash and MiMo V2.5 Pro pair 0.97+ means with waits under a minute a task at median. GLM 5.3 Flash matches them on score but takes over six times as long — fine for batch jobs, wrong for anything live.
Judgment calls on messy data (renewal risk with conflicting dates, churn-save eligibility): even the best model missed 3 of 40. Keep a human in the loop on high-stakes calls.
Cost discipline
The value tier (21 models at ≥0.93 mean) makes premium pricing hard to defend for routine operational work. Measure your own workload before paying a premium.
Reasoning-heavy pro models (the 5-pro class and successors) never beat the top open models on these tasks — their premium buys speed consistency and vendor SLAs, not accuracy on operational work.
Guardrails that matter regardless of model
Entity whitelist validation: 22 of the 119 models invented a name, account, or owner at least once across 4,684 graded answers. A deterministic alias check caught every case. It costs nothing to run.
Deterministic spot-audit: re-derive key numbers from source data on a schedule. Even the top scores hid subtle misses on the hardest tests.
"Cannot be determined" acceptance: reward models for saying the data cannot answer. The best models said so when appropriate; weaker ones guessed instead.
⚖️ These picks come from measured behavior on GTM work, not from vendor claims. Your data may shuffle the top tier. Test 20 to 40 of your own jobs before you spend.