← RevenueBench home
RevenueBench · Behavioral Clustering · 2026-09-27

Model Archetypes — How the Models Actually Behaved

118 models with ≥20 tests, clustered on behavioral dimensions — not just scores: per-category strengths, cost, latency, fabrication events, and math-vs-prose tilt. This is an acknowledgement of observed behavior, not a recommendation. Every model links to its full answer-by-answer record in the Response Reviewer.

Use these profiles conversationally: "these models perform like this." A Complete Analyst behaves like a senior operator; a Fast Communicator writes well but stumbles on rollforward math; Capable-but-Loose models need an entity-validation guardrail. Clustering is k-means (k=6, 20 seeds) over standardized category profiles + log-cost + log-latency + fabrication rate.

The Complete Analysts

39 models
elite-accuracyfastall-around-analyst
median mean 0.952median latency 56sstrongest marketingweakest data & crmfabrication events 0
Elite-accuracy, fast, all-around performers. High scores in every category with no weak flank; communication and reporting both strong. These models behave like a strong senior analyst: correct math, disciplined formatting, honest caveats.
ModelMeanLat
fireworks/ember-10.980125s
meta/muse-spark-1.10.97740s
moonshotai/kimi-k30.976102s
meta/muse-spark-1.30.97343s
unbiased/pareto0.97367s
google/gemini-3.8-flash0.970117s
anthropic/claude-opus-50.967107s
qwen/qwen3.8-max-09020.966160s
aion-labs/aion-3.50.966195s
z-ai/glm-5.3-prime0.966232s
meta/muse-spark-1.20.96462s
openai/gpt-5.5-pro0.961183s
z-ai/glm-5.30.960301s
anthropic/claude-opus-5.50.95863s
google/gemini-3.7-flash0.95783s
anthropic/claude-opus-4.50.95790s
sakana/fugu-ultra0.957410s
qwen/qwen3.8-max-prime0.953109s
openai/gpt-5.2-pro0.952181s
openai/gpt-6-astra0.95248s
xiaomi/mimo-v2.6-pro-ultraspeed0.95148s
anthropic/claude-opus-4.70.95062s
qwen/qwen3.8-2.4t-a95b0.948259s
openai/gpt-5.50.94537s
openai/gpt-5.20.944106s
openai/gpt-5.4-pro0.942256s
anthropic/claude-opus-4.80.94177s
openai/gpt-6-astra-pro0.94169s
openai/gpt-6-sol0.94154s
anthropic/claude-sonnet-50.940115s
openai/gpt-5.6-sol0.94039s
anthropic/claude-haiku-4.50.94059s
openai/gpt-5.6-sol-pro0.93753s
openai/gpt-5.10.935115s
openai/gpt-5-pro0.927355s
openai/gpt-6-sol-pro0.92176s
openai/gpt-5.6-terra0.92027s
thinkingmachines/inkling0.91647s
openai/gpt-5.6-terra-pro0.91541s

The Capable but Loose

37 models
elite-accuracybudgetprose-over-math10 fabrication-events
median mean 0.947median latency 74sstrongest marketingweakest reporting & analyticsfabrication events 10
Strong scores at budget prices, but this cluster carries fabrication events — models that invented a deal or company alias at least once. Correct on most work, but their output needs entity-validation guardrails before it touches production.
ModelMeanLat
deepseek/deepseek-pro-latest0.985177s
xiaomi/mimo-v2.6-pro0.984164s
z-ai/glm-5.3-flashx0.978115s
xiaomi/mimo-v2.5-pro0.974119s
qwen/qwen3.8-flash0.969111s
z-ai/glm-5.3-flash0.968401s
z-ai/glm-5.10.968178s
tencent/hy3-preview0.96882s
z-ai/glm-5v-turbo0.967154s
tencent/hy30.9671611s
z-ai/glm-5-turbo0.966141s
xiaomi/mimo-v2.6-flash0.962170s
qwen/qwen3.8-omni-flash0.962185s
deepseek/deepseek-v4-pro0.959328s
deepseek/deepseek-flash-latest0.956151s
meituan/longcat-2.00.955245s
moonshotai/kimi-k2.7-code0.9531145s
qwen/qwen3.8-27b0.953213s
deepseek/deepseek-v4-pro-08130.947110s
minimax/minimax-m2.10.94260s
stealth/space-bunny-alpha0.942143s
deepseek/deepseek-v4-flash-07310.93955s
poolside/laguna-s-2.10.939173s
openai/gpt-5.6-luna0.93641s
xiaomi/mimo-v2.50.93663s
deepseek/deepseek-v4-flash0.936151s
openai/gpt-5.4-mini0.93572s
openai/gpt-5.6-luna-pro0.93475s
stepfun/step-3.5-flash0.928166s
meta/muse-glimmer-30b0.92598s
openai/gpt-5.4-nano0.920117s
inclusionai/ling-3.0-flash0.91949s
minimax/minimax-m2.50.919594s
minimax/minimax-m20.91459s
upstage/solar-pro40.90954s
qwen/qwen3.7-flash0.90683s
openai/gpt-6-luna0.90468s

The Reliable Operators

23 models
midbudgetprose-over-math3 fabrication-events
median mean 0.880median latency 99sstrongest ops & maintenanceweakest data & crmfabrication events 3
Strong, budget-priced workhorses. Slightly below the elite tier, best on ops-maintenance and routine execution, weakest on data-crm edge cases. Behavior profile: steady, cheap, occasionally thin on the hardest multi-table reconciliation.
ModelMeanLat
nex-agi/nex-n2-mini0.9391445s
aion-labs/aion-3.5-mini0.926166s
stepfun/step-3.7-flash0.9191536s
openai/gpt-50.916131s
ibm-granite/granite-4.2-8b0.912310s
tencent/hy4-preview0.9061548s
aion-labs/aion-3.00.904913s
bytedance-seed/seed-2.0-mini0.900118s
openai/gpt-chat-latest0.899120s
openai/gpt-5-mini0.896120s
google/gemma-4-31b-it0.889516s
bytedance-seed/seed-1.60.880116s
minimax/minimax-m30.8782415s
aion-labs/aion-3.0-mini0.874194s
openai/gpt-6-luna-pro0.87479s
arcee-ai/trinity-large-thinking0.86974s
google/gemini-3-pro-image0.86866s
deepseek/deepseek-v3.20.863220s
cohere/command-a-plus0.86378s
poolside/laguna-xs-2.10.858623s
aion-labs/aion-2.00.853125s
thinkingmachines/inkling-small0.85158s
google/gemma-4-26b-a4b-it0.8491002s

The Fast Communicators

8 models
weakbudgetfastprose-over-math
median mean 0.734median latency 47sstrongest ops & maintenanceweakest reporting & analyticsfabrication events 0
Budget, fast, prose-oriented. Comfortable writing digests and emails, weak on the heavy reporting math (weighted forecasts, rollforwards). Behavior profile: good words, shaky arithmetic — fit for drafts, not for numbers.
ModelMeanLat
openai/gpt-5-nano0.790106s
openai/o4-mini0.77253s
mistralai/ministral-14b-25120.75565s
inclusionai/ling-3.0-flash-fin0.741727s
nvidia/nemotron-3-nano-30b-a3b0.728133s
openai/o30.71543s
meta-llama/llama-4-maverick0.689151s
meta-llama/llama-3.3-70b-instruct0.68375s

The Speed Specialists

8 models
weakbudgetfastprose-over-math9 fabrication-events
median mean 0.719median latency 27sstrongest ops & maintenanceweakest reporting & analyticsfabrication events 9
Budget models built for speed (median 27s) with below-mid accuracy and fabrication risk in-cluster. Trade correctness for latency; usable where a quick pass beats a right pass, with validation.
ModelMeanLat
mistralai/devstral-25120.79130s
inception/mercury-20.789746s
mistralai/mistral-medium-3-50.734873s
mistralai/ministral-3b-25120.73271s
mistralai/mistral-small-26030.7079s
amazon/nova-premier-v10.694102s
upstage/solar-pro-30.645182s
amazon/nova-pro-v10.61448s

The Fabricators

3 models
weakbudgetfastprose-over-math6 fabrication-events
median mean 0.814median latency 8sstrongest communicationweakest data & crmfabrication events 6
Small cluster where every member invented an entity at least once. Scores run low to middling with fabrication risk on top — verify every name and number before trusting output.
ModelMeanLat
inception/mercury-2.5-preview0.86723s
upstage/solar-mini40.81454s
nvidia/nemotron-3.5-lightning0.77622s

Cost for Value — The Dollar Map

106 full-coverage models
Where price and performance actually meet. The value frontier is the pareto set: no model is both cheaper and better than these.

THE VALUE FRONTIER — cheapest at every score level

ModelMean
deepseek/deepseek-pro-latest0.985
xiaomi/mimo-v2.5-pro0.974
qwen/qwen3.8-flash0.969
z-ai/glm-5.3-flash0.968
deepseek/deepseek-v4-flash-07310.939
inclusionai/ling-3.0-flash0.919
upstage/solar-pro40.909

VALUE KINGS — ≥0.93 mean, top efficiency (12 of 21 shown)

These models deliver top-decile scores at a fraction of premium spend. mimo-v2.5-pro beats the best big-lab scores on this bench (opus-5, gemini-3.8-flash). glm-5.3-flash, tencent/hy3, qwen3.8-flash land within half a point of those scores.
ModelMean
deepseek/deepseek-pro-latest0.985
z-ai/glm-5.3-flashx0.978
xiaomi/mimo-v2.5-pro0.974
qwen/qwen3.8-flash0.969
z-ai/glm-5.3-flash0.968
z-ai/glm-5.10.968
tencent/hy3-preview0.968
tencent/hy30.967
qwen/qwen3.8-omni-flash0.962
deepseek/deepseek-v4-pro0.959
deepseek/deepseek-flash-latest0.956
meituan/longcat-2.00.955

EXCELLENT VALUE — strong scores at modest prices

ModelMean
fireworks/ember-10.980
meta/muse-spark-1.10.977
moonshotai/kimi-k30.976
meta/muse-spark-1.30.973
google/gemini-3.8-flash0.970
qwen/qwen3.8-max-09020.966
meta/muse-spark-1.20.964
z-ai/glm-5.30.960
google/gemini-3.7-flash0.957
xiaomi/mimo-v2.6-pro-ultraspeed0.951

⚠ OVERPRICED — paid the premium, didn't cash it

High spend that the scores didn't justify. gpt-5.4-pro is the priciest run of the entire exercise, 16 of the 21 value kings outscore it. The premium tier's honest pitch is speed and vendor polish — on this bench, not accuracy.
ModelMean
openai/gpt-5.4-pro0.942
openai/gpt-5.5-pro0.961
openai/gpt-5.2-pro0.952
openai/gpt-5-pro0.927
openai/gpt-6-astra-pro0.941
Reading note: this is observed price-performance on these 40 tests, not a vendor verdict — list prices shift, and latency/SLA/enterprise features aren't priced here. The complete dollar map is in the Excel workbook and the value-map chart.