Model Archetypes — How the Models Actually Behaved
118 models with ≥20 tests, clustered on behavioral dimensions — not just scores: per-category strengths, cost, latency, fabrication events, and math-vs-prose tilt. This is an acknowledgement of observed behavior, not a recommendation. Every model links to its full answer-by-answer record in the Response Reviewer.
Use these profiles conversationally: "these models perform like this." A Complete Analyst behaves like a senior operator; a Fast Communicator writes well but stumbles on rollforward math; Capable-but-Loose models need an entity-validation guardrail. Clustering is k-means (k=6, 20 seeds) over standardized category profiles + log-cost + log-latency + fabrication rate.
The Complete Analysts
39 models
elite-accuracyfastall-around-analyst
median mean 0.952median latency 56sstrongest marketingweakest data & crmfabrication events 0
Elite-accuracy, fast, all-around performers. High scores in every category with no weak flank; communication and reporting both strong. These models behave like a strong senior analyst: correct math, disciplined formatting, honest caveats.
median mean 0.947median latency 74sstrongest marketingweakest reporting & analyticsfabrication events 10
Strong scores at budget prices, but this cluster carries fabrication events — models that invented a deal or company alias at least once. Correct on most work, but their output needs entity-validation guardrails before it touches production.
median mean 0.880median latency 99sstrongest ops & maintenanceweakest data & crmfabrication events 3
Strong, budget-priced workhorses. Slightly below the elite tier, best on ops-maintenance and routine execution, weakest on data-crm edge cases. Behavior profile: steady, cheap, occasionally thin on the hardest multi-table reconciliation.
median mean 0.734median latency 47sstrongest ops & maintenanceweakest reporting & analyticsfabrication events 0
Budget, fast, prose-oriented. Comfortable writing digests and emails, weak on the heavy reporting math (weighted forecasts, rollforwards). Behavior profile: good words, shaky arithmetic — fit for drafts, not for numbers.
median mean 0.719median latency 27sstrongest ops & maintenanceweakest reporting & analyticsfabrication events 9
Budget models built for speed (median 27s) with below-mid accuracy and fabrication risk in-cluster. Trade correctness for latency; usable where a quick pass beats a right pass, with validation.
median mean 0.814median latency 8sstrongest communicationweakest data & crmfabrication events 6
Small cluster where every member invented an entity at least once. Scores run low to middling with fabrication risk on top — verify every name and number before trusting output.
VALUE KINGS — ≥0.93 mean, top efficiency (12 of 21 shown)
These models deliver top-decile scores at a fraction of premium spend. mimo-v2.5-pro beats the best big-lab scores on this bench (opus-5, gemini-3.8-flash). glm-5.3-flash, tencent/hy3, qwen3.8-flash land within half a point of those scores.
High spend that the scores didn't justify. gpt-5.4-pro is the priciest run of the entire exercise, 16 of the 21 value kings outscore it. The premium tier's honest pitch is speed and vendor polish — on this bench, not accuracy.
Reading note: this is observed price-performance on these 40 tests, not a vendor verdict — list prices shift, and latency/SLA/enterprise features aren't priced here. The complete dollar map is in the Excel workbook and the value-map chart.