Last refreshed 2026-09-27 · 119 models · 40 tasks · graded by code, not by a model

Results: 119 models tested, 106 finished the full 40.

Every model below faced the same 40 revenue-operations jobs and the same fixed grading code. Ranked by mean score. 13 models stopped short of the full set; their rows sit outside the ranking with counts shown.

models tested
119
full coverage
106
partial runs
13
01 · TOP 15

Top 15 of 106 full-coverage models

Mean score across all 40 tasks. Each model name opens its graded answers in the reviewer.

#ModelMeanPerfectAvg latencyModalities
1deepseek/deepseek-pro-latest0.984637/40176.9stext
2fireworks/ember-10.979636/40124.8stextimage
3z-ai/glm-5.3-flashx0.978136/40114.6stextimagevideo
4meta/muse-spark-1.10.977335/4040.2stextimagefilevideoaudio
5moonshotai/kimi-k30.976236/40102.1stextimagevideo
6xiaomi/mimo-v2.5-pro0.973836/40119.1stext
7meta/muse-spark-1.30.973335/4043.1stextimagefilevideoaudio
8unbiased/pareto0.973336/4067.3stextimage
9google/gemini-3.8-flash frontier0.970434/40117.1stextimagefilevideoaudio
10qwen/qwen3.8-flash0.969235/40111.3stextimagevideo
11z-ai/glm-5.3-flash0.967935/40400.9stextimagevideo
12z-ai/glm-5.10.967933/40178.4stext
13tencent/hy3-preview0.967734/4082.2stext
14anthropic/claude-opus-5 frontier0.967135/40106.9stextimagefile
15tencent/hy30.967136/401611.3stext

Scores compare within this run only. Frontier baselines carry a frontier tag. Modalities in means what you can send the model: text, images, files, video, or audio. Modalities out means what it can send back: text for 118 of 119 models, plus images for gemini-3-pro-image. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what each model can handle in real work.

02 · FULL BOARD

All 106 full-coverage models

Ranked by mean score, best first. Model names link to graded answers in the reviewer.

#ModelMeanPerfectAvg latencyModalities
1deepseek/deepseek-pro-latest0.984637/40176.9stext
2fireworks/ember-10.979636/40124.8stextimage
3z-ai/glm-5.3-flashx0.978136/40114.6stextimagevideo
4meta/muse-spark-1.10.977335/4040.2stextimagefilevideoaudio
5moonshotai/kimi-k30.976236/40102.1stextimagevideo
6xiaomi/mimo-v2.5-pro0.973836/40119.1stext
7meta/muse-spark-1.30.973335/4043.1stextimagefilevideoaudio
8unbiased/pareto0.973336/4067.3stextimage
9google/gemini-3.8-flash frontier0.970434/40117.1stextimagefilevideoaudio
10qwen/qwen3.8-flash0.969235/40111.3stextimagevideo
11z-ai/glm-5.3-flash0.967935/40400.9stextimagevideo
12z-ai/glm-5.10.967933/40178.4stext
13tencent/hy3-preview0.967734/4082.2stext
14anthropic/claude-opus-5 frontier0.967135/40106.9stextimagefile
15tencent/hy30.967136/401611.3stext
16qwen/qwen3.8-max-09020.966235/40159.5stextimagevideo
17meta/muse-spark-1.20.964233/4062.3stextimagefilevideoaudio
18qwen/qwen3.8-omni-flash0.961733/40184.7stextimagevideoaudio
19openai/gpt-5.5-pro frontier0.960633/40183.4stextimagefile
20z-ai/glm-5.30.960433/40301.3stext
21deepseek/deepseek-v4-pro0.959233/40327.8stext
22anthropic/claude-opus-5.5 frontier0.957833/4063.2stextimagefile
23google/gemini-3.7-flash frontier0.957531/4083.5stextimagefilevideoaudio
24anthropic/claude-opus-4.5 frontier0.957132/4090.3stextimagefile
25sakana/fugu-ultra0.956734/40410.2stextimage
26deepseek/deepseek-flash-latest0.956332/40151.4stextimage
27meituan/longcat-2.00.955432/40245.4stext
28moonshotai/kimi-k2.7-code0.953131/401144.9stextimage
29qwen/qwen3.8-27b0.952932/40213.1stextimagevideo
30qwen/qwen3.8-max-prime0.952532/40108.8stextimagevideo
31openai/gpt-5.2-pro frontier0.952531/40180.7stextimagefile
32openai/gpt-6-astra frontier0.951734/4048.0stextimagefile
33xiaomi/mimo-v2.6-pro-ultraspeed0.950833/4048.3stextimagevideoaudio
34anthropic/claude-opus-4.7 frontier0.950033/4062.1stextimagefile
35qwen/qwen3.8-2.4t-a95b0.947532/40259.1stext
36deepseek/deepseek-v4-pro-08130.947132/40110.1stext
37openai/gpt-5.5 frontier0.945431/4036.9stextimagefile
38openai/gpt-5.2 frontier0.943929/40106.4stextimagefile
39minimax/minimax-m2.10.942531/4060.0stext
40openai/gpt-5.4-pro frontier0.942132/40256.4stextimagefile
41anthropic/claude-opus-4.8 frontier0.941429/4076.6stextimagefile
42openai/gpt-6-astra-pro frontier0.941232/4069.4stextimagefile
43openai/gpt-6-sol frontier0.941029/4054.3stextimagefile
44openai/gpt-5.6-sol frontier0.940229/4039.2stextimagefile
45anthropic/claude-sonnet-5 frontier0.940232/40115.1stextimagefile
46anthropic/claude-haiku-4.5 frontier0.939829/4058.8stextimagefile
47deepseek/deepseek-v4-flash-07310.939330/4055.2stext
48openai/gpt-5.6-sol-pro frontier0.937131/4053.2stextimagefile
49openai/gpt-5.6-luna frontier0.936330/4041.0stextimagefile
50xiaomi/mimo-v2.50.936331/4062.9stextimagevideoaudio
51deepseek/deepseek-v4-flash0.935830/40150.6stext
52openai/gpt-5.4-mini frontier0.935430/4072.3stextimagefile
53openai/gpt-5.1 frontier0.935330/40114.6stextimagefile
54openai/gpt-5.6-luna-pro frontier0.933529/4074.8stextimagefile
55stepfun/step-3.5-flash0.927729/40165.8stext
56openai/gpt-5-pro frontier0.927129/40354.8stextimagefile
57meta/muse-glimmer-30b0.924630/4097.8stextimage
58openai/gpt-6-sol-pro frontier0.921429/4076.3stextimagefile
59openai/gpt-5.4-nano frontier0.919726/40117.1stextimagefile
60openai/gpt-5.6-terra frontier0.919627/4026.9stextimagefile
61inclusionai/ling-3.0-flash0.919428/4048.6stext
62stepfun/step-3.7-flash0.919328/401536.5stextimagevideo
63minimax/minimax-m2.50.918727/40593.7stext
64openai/gpt-5 frontier0.915928/40131.1stextimagefile
65thinkingmachines/inkling0.915627/4047.2stextimageaudio
66openai/gpt-5.6-terra-pro frontier0.915026/4041.3stextimagefile
67minimax/minimax-m20.913728/4058.8stext
68upstage/solar-pro40.909229/4053.9stext
69tencent/hy4-preview0.906429/401548.2stext
70qwen/qwen3.7-flash0.905827/4083.5stextimagevideo
71openai/gpt-6-luna frontier0.904527/4068.0stextimagefile
72bytedance-seed/seed-2.0-mini0.899628/40118.1stextimagevideo
73openai/gpt-chat-latest frontier0.898827/40119.8stextimagefile
74openai/gpt-5-mini frontier0.895830/40120.1stextimagefile
75google/gemma-4-31b-it0.889224/40516.5stextimagevideo
76bytedance-seed/seed-1.60.880026/40116.1stextimagevideo
77minimax/minimax-m30.878126/402415.2stextimagevideo
78aion-labs/aion-3.0-mini0.874228/40194.3stext
79openai/gpt-6-luna-pro frontier0.873525/4078.7stextimagefile
80arcee-ai/trinity-large-thinking0.869226/4073.6stext
81google/gemini-3-pro-image frontier0.867526/4065.8stextimageout:image
82inception/mercury-2.5-preview0.866626/4022.9stext
83deepseek/deepseek-v3.20.863523/40220.0stext
84cohere/command-a-plus0.862724/4078.0stextimage
85poolside/laguna-xs-2.10.857920/40623.0stext
86aion-labs/aion-2.00.853326/40124.8stext
87thinkingmachines/inkling-small0.851323/4057.6stextimageaudio
88google/gemma-4-26b-a4b-it0.849220/401001.6stextimagevideo
89upstage/solar-mini40.813720/4053.7stext
90mistralai/devstral-25120.790816/4030.3stextfile
91openai/gpt-5-nano frontier0.789617/40106.5stextimagefile
92inception/mercury-20.788619/40745.8stext
93nvidia/nemotron-3.5-lightning0.776015/4022.3stext
94openai/o4-mini frontier0.772517/4053.2stextimagefile
95mistralai/ministral-14b-25120.754619/4064.8stextimage
96inclusionai/ling-3.0-flash-fin0.740617/40726.7stext
97mistralai/mistral-medium-3-50.734314/40873.0stextimagefile
98mistralai/ministral-3b-25120.732115/4071.5stextimage
99nvidia/nemotron-3-nano-30b-a3b0.727515/40133.4stext
100openai/o3 frontier0.715017/4043.1stextimagefile
101mistralai/mistral-small-26030.706810/408.7stextimage
102amazon/nova-premier-v1 frontier0.693916/40101.5stextimage
103meta-llama/llama-4-maverick0.689314/40151.3stextimage
104meta-llama/llama-3.3-70b-instruct0.682916/4075.0stext
105upstage/solar-pro-30.645013/40182.4stext
106amazon/nova-pro-v1 frontier0.61389/4047.7stextimage
03 · PARTIALS

13 partial runs, disclosed

These models stopped short of the full 40 tasks, mostly on slow responses that timed out. Ungraded tasks are excluded, never scored as zero, so these rows stay out of the ranking.

ModelTasks runMeanPerfectSpendStatusModalities
xiaomi/mimo-v2.6-pro38/400.983835partialtextimagevideoaudio
z-ai/glm-5v-turbo36/400.967131partialtextimagevideo
aion-labs/aion-3.533/400.966228partialtext
z-ai/glm-5-turbo31/400.966127partialtext
z-ai/glm-5.3-prime37/400.965931partialtext
xiaomi/mimo-v2.6-flash36/400.962031partialtextimagevideoaudio
stealth/space-bunny-alpha38/400.942128partialtextimagevideo
poolside/laguna-s-2.139/400.938933partialtext
nex-agi/nex-n2-mini28/400.938722partialtext
nex-agi/nex-n2-pro14/400.928610partialtext
aion-labs/aion-3.5-mini36/400.926128partialtext
ibm-granite/granite-4.2-8b39/400.911524partialtext
aion-labs/aion-3.039/400.904332partialtext
04 · CELL MATRIX

Every model, every task

The cell matrix shows one score per model per task: 106 ranked rows plus 13 partial rows across all 40 tasks, each cell linking to the graded answer.

Open the cell matrix

05 · DOWNLOADS

Take the data

Same run, two formats. Numbers match the boards above.

Spreadsheet

Results workbook

Leaderboard, full matrix, per-test difficulty, model info. Four sheets.

Download xlsx

Answers

Answer reviewer

Side-by-side: what each model wrote against the computed answer.

Open reviewer

About me

Amani Phipps, Revenue Architect at Bonusly

I saw this problem in my work. Others shared it too. So I helped build this.

Happy to talk if you want to go deeper.

Connect on LinkedIn →