A model that cannot see an image will never pass a screenshot test, no matter how it ranks. Before comparing scores, check what a model can actually take in. Pick a group below.
Every model has an input side and an output side. The input side is what you can hand it: text (a prompt), image (a screenshot or photo), file (a PDF or document attachment), video, and audio. The output side is what it can produce — every model in this run writes text; a few also generate images.
This is not a quality feature. It is a capability feature. A text-only model at the top of the leaderboard is still the wrong tool for "here's a screenshot of the bug." The benchmark tests operational GTM work in text, so every model here could run it — but the moment your job involves looking at something, the field narrows.
Why the word modality. It means data type. Text is one. Image is another. A model's modality is its senses. Text plus image reads a screenshot directly. Text only cannot.
Background: what multimodal models are. Two paths exist. Native reads image and text together. Pipeline turns the image into text first. Your IDE usually handles the pipeline for you. That holds until the job needs the look itself.
Worth knowing with open models: vision varies by endpoint. Some DeepSeek endpoints are text only, while Flash adds it (see the vision guide). If you plan real work on open models, check the endpoint takes what you send.
Grouped by input modalities from the live OpenRouter catalog. A model appears in exactly one group, ordered by input range. Pick a group:
The widest input range in the run: text, images, files, and in many cases video or audio. Best fit for agent roles that meet people where they are — Slack, email, shared drives.
| Model | Mean | Perfect | Reads |
|---|---|---|---|
| xiaomi/mimo-v2.6-pro | 0.984 | 35/38 | text, image, video, audio |
| z-ai/glm-5.3-flashx | 0.978 | 36/40 | text, image, video |
| meta/muse-spark-1.1 | 0.977 | 35/40 | text, image, file, video, audio |
| moonshotai/kimi-k3 | 0.976 | 36/40 | text, image, video |
| meta/muse-spark-1.3 | 0.973 | 35/40 | text, image, file, video, audio |
| google/gemini-3.8-flash | 0.970 | 34/40 | text, image, file, video, audio |
| qwen/qwen3.8-flash | 0.969 | 35/40 | text, image, video |
| z-ai/glm-5.3-flash | 0.968 | 35/40 | text, image, video |
Ranked by mean score in the 40-task run. Modality data comes from each model's published architecture, re-checked at every monthly refresh. n/40 means the model ran every task; partial runs are labeled in the full matrix.
Everything a text model does, plus screenshots and photos. Most also read attached files. This is the floor for any job where a person might paste in what they are looking at.
| Model | Mean | Perfect | Reads |
|---|---|---|---|
| fireworks/ember-1 | 0.980 | 36/40 | text, image |
| unbiased/pareto | 0.973 | 36/40 | text, image |
| anthropic/claude-opus-5 | 0.967 | 35/40 | text, image, file |
| openai/gpt-5.5-pro | 0.961 | 33/40 | text, image, file |
| anthropic/claude-opus-5.5 | 0.958 | 33/40 | text, image, file |
| anthropic/claude-opus-4.5 | 0.957 | 32/40 | text, image, file |
| sakana/fugu-ultra | 0.957 | 34/40 | text, image |
| deepseek/deepseek-flash-latest | 0.956 | 32/40 | text, image |
Ranked by mean score in the 40-task run. Modality data comes from each model's published architecture, re-checked at every monthly refresh. n/40 means the model ran every task; partial runs are labeled in the full matrix.
Text in, text out. No screenshots, no media. One of them also reads file attachments. Ask one to look at an image and it simply cannot.
| Model | Mean | Perfect | Reads |
|---|---|---|---|
| deepseek/deepseek-pro-latest | 0.985 | 37/40 | text |
| xiaomi/mimo-v2.5-pro | 0.974 | 36/40 | text |
| z-ai/glm-5.1 | 0.968 | 33/40 | text |
| tencent/hy3-preview | 0.968 | 34/40 | text |
| tencent/hy3 | 0.967 | 36/40 | text |
| aion-labs/aion-3.5 | 0.966 | 28/33 | text |
| z-ai/glm-5-turbo | 0.966 | 27/31 | text |
| z-ai/glm-5.3-prime | 0.966 | 31/37 | text |
Ranked by mean score in the 40-task run. Modality data comes from each model's published architecture, re-checked at every monthly refresh. n/40 means the model ran every task; partial runs are labeled in the full matrix.