Every model below faced the same 40 revenue-operations jobs and the same fixed grading code. Ranked by mean score. 13 models stopped short of the full set; their rows sit outside the ranking with counts shown.
Mean score across all 40 tasks. Each model name opens its graded answers in the reviewer.
| # | Model | Mean | Perfect | Avg latency | Modalities |
|---|---|---|---|---|---|
| 1 | deepseek/deepseek-pro-latest | 0.9846 | 37/40 | 176.9s | text |
| 2 | fireworks/ember-1 | 0.9796 | 36/40 | 124.8s | textimage |
| 3 | z-ai/glm-5.3-flashx | 0.9781 | 36/40 | 114.6s | textimagevideo |
| 4 | meta/muse-spark-1.1 | 0.9773 | 35/40 | 40.2s | textimagefilevideoaudio |
| 5 | moonshotai/kimi-k3 | 0.9762 | 36/40 | 102.1s | textimagevideo |
| 6 | xiaomi/mimo-v2.5-pro | 0.9738 | 36/40 | 119.1s | text |
| 7 | meta/muse-spark-1.3 | 0.9733 | 35/40 | 43.1s | textimagefilevideoaudio |
| 8 | unbiased/pareto | 0.9733 | 36/40 | 67.3s | textimage |
| 9 | google/gemini-3.8-flash frontier | 0.9704 | 34/40 | 117.1s | textimagefilevideoaudio |
| 10 | qwen/qwen3.8-flash | 0.9692 | 35/40 | 111.3s | textimagevideo |
| 11 | z-ai/glm-5.3-flash | 0.9679 | 35/40 | 400.9s | textimagevideo |
| 12 | z-ai/glm-5.1 | 0.9679 | 33/40 | 178.4s | text |
| 13 | tencent/hy3-preview | 0.9677 | 34/40 | 82.2s | text |
| 14 | anthropic/claude-opus-5 frontier | 0.9671 | 35/40 | 106.9s | textimagefile |
| 15 | tencent/hy3 | 0.9671 | 36/40 | 1611.3s | text |
Scores compare within this run only. Frontier baselines carry a frontier tag. Modalities in means what you can send the model: text, images, files, video, or audio. Modalities out means what it can send back: text for 118 of 119 models, plus images for gemini-3-pro-image. Bench tasks feed text extracts, so modalities do not change the score. They are here so you know what each model can handle in real work.
Ranked by mean score, best first. Model names link to graded answers in the reviewer.
| # | Model | Mean | Perfect | Avg latency | Modalities |
|---|---|---|---|---|---|
| 1 | deepseek/deepseek-pro-latest | 0.9846 | 37/40 | 176.9s | text |
| 2 | fireworks/ember-1 | 0.9796 | 36/40 | 124.8s | textimage |
| 3 | z-ai/glm-5.3-flashx | 0.9781 | 36/40 | 114.6s | textimagevideo |
| 4 | meta/muse-spark-1.1 | 0.9773 | 35/40 | 40.2s | textimagefilevideoaudio |
| 5 | moonshotai/kimi-k3 | 0.9762 | 36/40 | 102.1s | textimagevideo |
| 6 | xiaomi/mimo-v2.5-pro | 0.9738 | 36/40 | 119.1s | text |
| 7 | meta/muse-spark-1.3 | 0.9733 | 35/40 | 43.1s | textimagefilevideoaudio |
| 8 | unbiased/pareto | 0.9733 | 36/40 | 67.3s | textimage |
| 9 | google/gemini-3.8-flash frontier | 0.9704 | 34/40 | 117.1s | textimagefilevideoaudio |
| 10 | qwen/qwen3.8-flash | 0.9692 | 35/40 | 111.3s | textimagevideo |
| 11 | z-ai/glm-5.3-flash | 0.9679 | 35/40 | 400.9s | textimagevideo |
| 12 | z-ai/glm-5.1 | 0.9679 | 33/40 | 178.4s | text |
| 13 | tencent/hy3-preview | 0.9677 | 34/40 | 82.2s | text |
| 14 | anthropic/claude-opus-5 frontier | 0.9671 | 35/40 | 106.9s | textimagefile |
| 15 | tencent/hy3 | 0.9671 | 36/40 | 1611.3s | text |
| 16 | qwen/qwen3.8-max-0902 | 0.9662 | 35/40 | 159.5s | textimagevideo |
| 17 | meta/muse-spark-1.2 | 0.9642 | 33/40 | 62.3s | textimagefilevideoaudio |
| 18 | qwen/qwen3.8-omni-flash | 0.9617 | 33/40 | 184.7s | textimagevideoaudio |
| 19 | openai/gpt-5.5-pro frontier | 0.9606 | 33/40 | 183.4s | textimagefile |
| 20 | z-ai/glm-5.3 | 0.9604 | 33/40 | 301.3s | text |
| 21 | deepseek/deepseek-v4-pro | 0.9592 | 33/40 | 327.8s | text |
| 22 | anthropic/claude-opus-5.5 frontier | 0.9578 | 33/40 | 63.2s | textimagefile |
| 23 | google/gemini-3.7-flash frontier | 0.9575 | 31/40 | 83.5s | textimagefilevideoaudio |
| 24 | anthropic/claude-opus-4.5 frontier | 0.9571 | 32/40 | 90.3s | textimagefile |
| 25 | sakana/fugu-ultra | 0.9567 | 34/40 | 410.2s | textimage |
| 26 | deepseek/deepseek-flash-latest | 0.9563 | 32/40 | 151.4s | textimage |
| 27 | meituan/longcat-2.0 | 0.9554 | 32/40 | 245.4s | text |
| 28 | moonshotai/kimi-k2.7-code | 0.9531 | 31/40 | 1144.9s | textimage |
| 29 | qwen/qwen3.8-27b | 0.9529 | 32/40 | 213.1s | textimagevideo |
| 30 | qwen/qwen3.8-max-prime | 0.9525 | 32/40 | 108.8s | textimagevideo |
| 31 | openai/gpt-5.2-pro frontier | 0.9525 | 31/40 | 180.7s | textimagefile |
| 32 | openai/gpt-6-astra frontier | 0.9517 | 34/40 | 48.0s | textimagefile |
| 33 | xiaomi/mimo-v2.6-pro-ultraspeed | 0.9508 | 33/40 | 48.3s | textimagevideoaudio |
| 34 | anthropic/claude-opus-4.7 frontier | 0.9500 | 33/40 | 62.1s | textimagefile |
| 35 | qwen/qwen3.8-2.4t-a95b | 0.9475 | 32/40 | 259.1s | text |
| 36 | deepseek/deepseek-v4-pro-0813 | 0.9471 | 32/40 | 110.1s | text |
| 37 | openai/gpt-5.5 frontier | 0.9454 | 31/40 | 36.9s | textimagefile |
| 38 | openai/gpt-5.2 frontier | 0.9439 | 29/40 | 106.4s | textimagefile |
| 39 | minimax/minimax-m2.1 | 0.9425 | 31/40 | 60.0s | text |
| 40 | openai/gpt-5.4-pro frontier | 0.9421 | 32/40 | 256.4s | textimagefile |
| 41 | anthropic/claude-opus-4.8 frontier | 0.9414 | 29/40 | 76.6s | textimagefile |
| 42 | openai/gpt-6-astra-pro frontier | 0.9412 | 32/40 | 69.4s | textimagefile |
| 43 | openai/gpt-6-sol frontier | 0.9410 | 29/40 | 54.3s | textimagefile |
| 44 | openai/gpt-5.6-sol frontier | 0.9402 | 29/40 | 39.2s | textimagefile |
| 45 | anthropic/claude-sonnet-5 frontier | 0.9402 | 32/40 | 115.1s | textimagefile |
| 46 | anthropic/claude-haiku-4.5 frontier | 0.9398 | 29/40 | 58.8s | textimagefile |
| 47 | deepseek/deepseek-v4-flash-0731 | 0.9393 | 30/40 | 55.2s | text |
| 48 | openai/gpt-5.6-sol-pro frontier | 0.9371 | 31/40 | 53.2s | textimagefile |
| 49 | openai/gpt-5.6-luna frontier | 0.9363 | 30/40 | 41.0s | textimagefile |
| 50 | xiaomi/mimo-v2.5 | 0.9363 | 31/40 | 62.9s | textimagevideoaudio |
| 51 | deepseek/deepseek-v4-flash | 0.9358 | 30/40 | 150.6s | text |
| 52 | openai/gpt-5.4-mini frontier | 0.9354 | 30/40 | 72.3s | textimagefile |
| 53 | openai/gpt-5.1 frontier | 0.9353 | 30/40 | 114.6s | textimagefile |
| 54 | openai/gpt-5.6-luna-pro frontier | 0.9335 | 29/40 | 74.8s | textimagefile |
| 55 | stepfun/step-3.5-flash | 0.9277 | 29/40 | 165.8s | text |
| 56 | openai/gpt-5-pro frontier | 0.9271 | 29/40 | 354.8s | textimagefile |
| 57 | meta/muse-glimmer-30b | 0.9246 | 30/40 | 97.8s | textimage |
| 58 | openai/gpt-6-sol-pro frontier | 0.9214 | 29/40 | 76.3s | textimagefile |
| 59 | openai/gpt-5.4-nano frontier | 0.9197 | 26/40 | 117.1s | textimagefile |
| 60 | openai/gpt-5.6-terra frontier | 0.9196 | 27/40 | 26.9s | textimagefile |
| 61 | inclusionai/ling-3.0-flash | 0.9194 | 28/40 | 48.6s | text |
| 62 | stepfun/step-3.7-flash | 0.9193 | 28/40 | 1536.5s | textimagevideo |
| 63 | minimax/minimax-m2.5 | 0.9187 | 27/40 | 593.7s | text |
| 64 | openai/gpt-5 frontier | 0.9159 | 28/40 | 131.1s | textimagefile |
| 65 | thinkingmachines/inkling | 0.9156 | 27/40 | 47.2s | textimageaudio |
| 66 | openai/gpt-5.6-terra-pro frontier | 0.9150 | 26/40 | 41.3s | textimagefile |
| 67 | minimax/minimax-m2 | 0.9137 | 28/40 | 58.8s | text |
| 68 | upstage/solar-pro4 | 0.9092 | 29/40 | 53.9s | text |
| 69 | tencent/hy4-preview | 0.9064 | 29/40 | 1548.2s | text |
| 70 | qwen/qwen3.7-flash | 0.9058 | 27/40 | 83.5s | textimagevideo |
| 71 | openai/gpt-6-luna frontier | 0.9045 | 27/40 | 68.0s | textimagefile |
| 72 | bytedance-seed/seed-2.0-mini | 0.8996 | 28/40 | 118.1s | textimagevideo |
| 73 | openai/gpt-chat-latest frontier | 0.8988 | 27/40 | 119.8s | textimagefile |
| 74 | openai/gpt-5-mini frontier | 0.8958 | 30/40 | 120.1s | textimagefile |
| 75 | google/gemma-4-31b-it | 0.8892 | 24/40 | 516.5s | textimagevideo |
| 76 | bytedance-seed/seed-1.6 | 0.8800 | 26/40 | 116.1s | textimagevideo |
| 77 | minimax/minimax-m3 | 0.8781 | 26/40 | 2415.2s | textimagevideo |
| 78 | aion-labs/aion-3.0-mini | 0.8742 | 28/40 | 194.3s | text |
| 79 | openai/gpt-6-luna-pro frontier | 0.8735 | 25/40 | 78.7s | textimagefile |
| 80 | arcee-ai/trinity-large-thinking | 0.8692 | 26/40 | 73.6s | text |
| 81 | google/gemini-3-pro-image frontier | 0.8675 | 26/40 | 65.8s | textimageout:image |
| 82 | inception/mercury-2.5-preview | 0.8666 | 26/40 | 22.9s | text |
| 83 | deepseek/deepseek-v3.2 | 0.8635 | 23/40 | 220.0s | text |
| 84 | cohere/command-a-plus | 0.8627 | 24/40 | 78.0s | textimage |
| 85 | poolside/laguna-xs-2.1 | 0.8579 | 20/40 | 623.0s | text |
| 86 | aion-labs/aion-2.0 | 0.8533 | 26/40 | 124.8s | text |
| 87 | thinkingmachines/inkling-small | 0.8513 | 23/40 | 57.6s | textimageaudio |
| 88 | google/gemma-4-26b-a4b-it | 0.8492 | 20/40 | 1001.6s | textimagevideo |
| 89 | upstage/solar-mini4 | 0.8137 | 20/40 | 53.7s | text |
| 90 | mistralai/devstral-2512 | 0.7908 | 16/40 | 30.3s | textfile |
| 91 | openai/gpt-5-nano frontier | 0.7896 | 17/40 | 106.5s | textimagefile |
| 92 | inception/mercury-2 | 0.7886 | 19/40 | 745.8s | text |
| 93 | nvidia/nemotron-3.5-lightning | 0.7760 | 15/40 | 22.3s | text |
| 94 | openai/o4-mini frontier | 0.7725 | 17/40 | 53.2s | textimagefile |
| 95 | mistralai/ministral-14b-2512 | 0.7546 | 19/40 | 64.8s | textimage |
| 96 | inclusionai/ling-3.0-flash-fin | 0.7406 | 17/40 | 726.7s | text |
| 97 | mistralai/mistral-medium-3-5 | 0.7343 | 14/40 | 873.0s | textimagefile |
| 98 | mistralai/ministral-3b-2512 | 0.7321 | 15/40 | 71.5s | textimage |
| 99 | nvidia/nemotron-3-nano-30b-a3b | 0.7275 | 15/40 | 133.4s | text |
| 100 | openai/o3 frontier | 0.7150 | 17/40 | 43.1s | textimagefile |
| 101 | mistralai/mistral-small-2603 | 0.7068 | 10/40 | 8.7s | textimage |
| 102 | amazon/nova-premier-v1 frontier | 0.6939 | 16/40 | 101.5s | textimage |
| 103 | meta-llama/llama-4-maverick | 0.6893 | 14/40 | 151.3s | textimage |
| 104 | meta-llama/llama-3.3-70b-instruct | 0.6829 | 16/40 | 75.0s | text |
| 105 | upstage/solar-pro-3 | 0.6450 | 13/40 | 182.4s | text |
| 106 | amazon/nova-pro-v1 frontier | 0.6138 | 9/40 | 47.7s | textimage |
These models stopped short of the full 40 tasks, mostly on slow responses that timed out. Ungraded tasks are excluded, never scored as zero, so these rows stay out of the ranking.
| Model | Tasks run | Mean | Perfect | Spend | Status | Modalities |
|---|---|---|---|---|---|---|
| xiaomi/mimo-v2.6-pro | 38/40 | 0.9838 | 35 | partial | textimagevideoaudio | |
| z-ai/glm-5v-turbo | 36/40 | 0.9671 | 31 | partial | textimagevideo | |
| aion-labs/aion-3.5 | 33/40 | 0.9662 | 28 | partial | text | |
| z-ai/glm-5-turbo | 31/40 | 0.9661 | 27 | partial | text | |
| z-ai/glm-5.3-prime | 37/40 | 0.9659 | 31 | partial | text | |
| xiaomi/mimo-v2.6-flash | 36/40 | 0.9620 | 31 | partial | textimagevideoaudio | |
| stealth/space-bunny-alpha | 38/40 | 0.9421 | 28 | partial | textimagevideo | |
| poolside/laguna-s-2.1 | 39/40 | 0.9389 | 33 | partial | text | |
| nex-agi/nex-n2-mini | 28/40 | 0.9387 | 22 | partial | text | |
| nex-agi/nex-n2-pro | 14/40 | 0.9286 | 10 | partial | text | |
| aion-labs/aion-3.5-mini | 36/40 | 0.9261 | 28 | partial | text | |
| ibm-granite/granite-4.2-8b | 39/40 | 0.9115 | 24 | partial | text | |
| aion-labs/aion-3.0 | 39/40 | 0.9043 | 32 | partial | text |
The cell matrix shows one score per model per task: 106 ranked rows plus 13 partial rows across all 40 tasks, each cell linking to the graded answer.
Same run, two formats. Numbers match the boards above.
Leaderboard, full matrix, per-test difficulty, model info. Four sheets.
I saw this problem in my work. Others shared it too. So I helped build this.
Happy to talk if you want to go deeper.