Modality data refreshed from the live model catalog · 119 models

Modality first. Score second.

A model that cannot see an image will never pass a screenshot test, no matter how it ranks. Before comparing scores, check what a model can actually take in. Pick a group below.

01 · WHAT MODALITIES ARE

What a model can take in, and what it can put out.

Every model has an input side and an output side. The input side is what you can hand it: text (a prompt), image (a screenshot or photo), file (a PDF or document attachment), video, and audio. The output side is what it can produce — every model in this run writes text; a few also generate images.

This is not a quality feature. It is a capability feature. A text-only model at the top of the leaderboard is still the wrong tool for "here's a screenshot of the bug." The benchmark tests operational GTM work in text, so every model here could run it — but the moment your job involves looking at something, the field narrows.

Why the word modality. It means data type. Text is one. Image is another. A model's modality is its senses. Text plus image reads a screenshot directly. Text only cannot.

Background: what multimodal models are. Two paths exist. Native reads image and text together. Pipeline turns the image into text first. Your IDE usually handles the pipeline for you. That holds until the job needs the look itself.

Worth knowing with open models: vision varies by endpoint. Some DeepSeek endpoints are text only, while Flash adds it (see the vision guide). If you plan real work on open models, check the endpoint takes what you send.

02 · THE THREE GROUPS

119 models, three modality archetypes.

Grouped by input modalities from the live OpenRouter catalog. A model appears in exactly one group, ordered by input range. Pick a group:

Media Generalists

The widest input range in the run: text, images, files, and in many cases video or audio. Best fit for agent roles that meet people where they are — Slack, email, shared drives.

28 of 119 modelsmedian mean 0.955text · image · file · video · audio
ModelMeanPerfectReads
xiaomi/mimo-v2.6-pro0.98435/38text, image, video, audio
z-ai/glm-5.3-flashx0.97836/40text, image, video
meta/muse-spark-1.10.97735/40text, image, file, video, audio
moonshotai/kimi-k30.97636/40text, image, video
meta/muse-spark-1.30.97335/40text, image, file, video, audio
google/gemini-3.8-flash0.97034/40text, image, file, video, audio
qwen/qwen3.8-flash0.96935/40text, image, video
z-ai/glm-5.3-flash0.96835/40text, image, video

Ranked by mean score in the 40-task run. Modality data comes from each model's published architecture, re-checked at every monthly refresh. n/40 means the model ran every task; partial runs are labeled in the full matrix.

Image Readers

Everything a text model does, plus screenshots and photos. Most also read attached files. This is the floor for any job where a person might paste in what they are looking at.

49 of 119 modelsmedian mean 0.935text · image · file
ModelMeanPerfectReads
fireworks/ember-10.98036/40text, image
unbiased/pareto0.97336/40text, image
anthropic/claude-opus-50.96735/40text, image, file
openai/gpt-5.5-pro0.96133/40text, image, file
anthropic/claude-opus-5.50.95833/40text, image, file
anthropic/claude-opus-4.50.95732/40text, image, file
sakana/fugu-ultra0.95734/40text, image
deepseek/deepseek-flash-latest0.95632/40text, image

Ranked by mean score in the 40-task run. Modality data comes from each model's published architecture, re-checked at every monthly refresh. n/40 means the model ran every task; partial runs are labeled in the full matrix.

Text Specialists

Text in, text out. No screenshots, no media. One of them also reads file attachments. Ask one to look at an image and it simply cannot.

42 of 119 modelsmedian mean 0.923text
ModelMeanPerfectReads
deepseek/deepseek-pro-latest0.98537/40text
xiaomi/mimo-v2.5-pro0.97436/40text
z-ai/glm-5.10.96833/40text
tencent/hy3-preview0.96834/40text
tencent/hy30.96736/40text
aion-labs/aion-3.50.96628/33text
z-ai/glm-5-turbo0.96627/31text
z-ai/glm-5.3-prime0.96631/37text

Ranked by mean score in the 40-task run. Modality data comes from each model's published architecture, re-checked at every monthly refresh. n/40 means the model ran every task; partial runs are labeled in the full matrix.