MMMU-Pro

    🏆 Leaderboard

    As of August 20, 2026, Gemini 3.5 Flash is #1 for MMMU-Pro at 83.6%. Ranked by MMMU-Pro: a harder multimodal exam than MMMU. 54 models in this index have a published MMMU-Pro score. MMMU-Pro leaderboard with live API prices. MMMU-Pro — a harder subset of MMMU with more complex multimodal reasoning.

    Updated August 20, 2026282 models33 providers
    Gemini 3.5 Flash
    Google · Proprietary
    83.6%$10.50
    GPT-5.5
    OpenAI · Proprietary
    83.2%$35.00
    GPT-5.6 Sol
    OpenAI · Proprietary
    83%$35.00
    Qwen3.8 MaxOSS
    Qwen · Open Source
    82.3%$8.00
    Seed 2.1 Turbo
    ByteDance · Proprietary
    82.2%$3.00
    Kimi K3OSS
    Moonshot AI · Open Source
    81.6%$18.00
    GPT-5.4
    OpenAI · Proprietary
    81.2%$17.50
    Gemini 3 Flash
    Google · Proprietary
    81.2%$3.50
    Gemini 3 Pro
    Google · Proprietary
    81%$14.00
    GPT-5.6 Terra
    OpenAI · Proprietary
    80.7%$14.00
    Gemini 3.1 Pro
    Google · Proprietary
    80.5%$17.50
    Kimi K2.6OSS
    Moonshot AI · Open Source
    80.1%$4.93
    GPT-5.2
    OpenAI · Proprietary
    79.5%$15.75
    Qwen3.7-Plus
    Qwen · Proprietary
    79%$1.60
    Qwen3.6 Plus
    Qwen · Proprietary
    78.8%$3.50
    Kimi K2.5OSS
    Moonshot AI · Open Source
    78.5%$3.68
    GPT-5.6 Luna
    OpenAI · Proprietary
    78.4%$1.40
    GPT-5
    OpenAI · Proprietary
    78.4%$11.25
    MiniMax M3OSS
    MiniMax · Open Source
    78.1%$1.50
    MiMo-V2.5OSS
    Xiaomi · Open Source
    77.9%$0.50
    Claude Opus 4.6
    Anthropic · Proprietary
    77.3%$30.00
    Qwen3.5-122B-A10BOSS
    Qwen · Open Source
    76.9%$3.60
    Gemma 4 31BOSS
    Google · Open Source
    76.9%$0.54
    Gemini 3.1 Flash-Lite
    Google · Proprietary
    76.8%$1.75
    GPT-5.4 mini
    OpenAI · Proprietary
    76.6%$5.25
    o3
    OpenAI · Proprietary
    76.4%$10.00
    GPT-5.5 Instant
    OpenAI · Proprietary
    76%$35.00
    Qwen3.6-27BOSS
    Qwen · Open Source
    75.8%$4.20
    Claude Sonnet 4.6
    Anthropic · Proprietary
    75.6%$18.00
    Qwen3.6-35B-A3BOSS
    Qwen · Open Source · via OpenRouter
    75.3%$1.14
    Qwen3.5-35B-A3BOSS
    Qwen · Open Source
    75.1%$2.25
    Qwen3.5-27BOSS
    Qwen · Open Source
    75%$2.70
    Muse Glimmer-30BOSS
    Meta · Open Source
    74%$1.85
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    74%$1.50
    Gemma 4 26B-A4BOSS
    Google · Open Source
    73.8%$0.53
    Qwen3 VL 235B A22B ThinkingOSS
    Qwen · Open Source
    69.3%$3.94
    Qwen3 VL 235B A22B InstructOSS
    Qwen · Open Source
    68.1%$1.79
    GPT-5.4 nano
    OpenAI · Proprietary
    66.1%$1.45
    Qwen3 VL 32B InstructOSS
    Qwen · Open Source · via OpenRouter
    65.3%$0.52
    Qwen3 VL 30B A3B ThinkingOSS
    Qwen · Open Source
    63%$1.20
    Command A+OSS
    Cohere · Open Source
    63%$12.50
    Nova 2 Lite
    Amazon · Proprietary
    61.8%$2.80
    Qwen3 VL 8B ThinkingOSS
    Qwen · Open Source
    60.4%$2.27
    Qwen3 VL 30B A3B InstructOSS
    Qwen · Open Source
    60.4%$0.90
    Mistral Small 4OSS
    Mistral · Open Source · via Mistral AI
    60%$0.75
    GPT-4o
    OpenAI · Proprietary
    59.9%$12.50
    Llama 4 MaverickOSS
    Meta · Open Source
    59.6%$1.02
    Qwen3 VL 4B ThinkingOSS
    Qwen · Open Source
    57%$1.10
    Qwen3 VL 8B InstructOSS
    Qwen · Open Source
    55.9%$0.58
    Qwen3 VL 4B InstructOSS
    Qwen · Open Source
    53.2%$0.70
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads MMMU-Pro right now?

    As of August 20, 2026, Gemini 3.5 Flash by Google is #1 for MMMU-Pro at 83.6%. Ranked by MMMU-Pro: a harder multimodal exam than MMMU. This board also tracks MMMU-Pro. Next on the same board: GPT-5.5 and GPT-5.6 Sol. This mmmu-pro leaderboard ranks models by MMMU-Pro. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for mmmu-pro. Ranked by MMMU-Pro: a harder multimodal exam than MMMU. Input and output are dollars per million tokens.
    RankModelMMMU-ProInput /MOutput /M
    1Gemini 3.5 Flash83.6%$1.50$9.00
    2GPT-5.583.2%$5.00$30.00
    3GPT-5.6 Sol83%$5.00$30.00
    4Qwen3.8 Max82.3%$2.00$6.00
    5Seed 2.1 Turbo82.2%$0.50$2.50
    6Kimi K381.6%$3.00$15.00
    7GPT-5.481.2%$2.50$15.00
    8Gemini 3 Flash81.2%$0.50$3.00

    MMMU-Pro FAQ

    Who ranks #1 on the MMMU-Pro leaderboard?

    As of August 20, 2026, Gemini 3.5 Flash by Google ranks #1 on MMMU-Pro at 83.6%. API pricing is $1.50/M input and $9.00/M output.

    What are the top models on MMMU-Pro?

    The current MMMU-Pro ranking as of August 20, 2026 is 1. Gemini 3.5 Flash at 83.6%; 2. GPT-5.5 at 83.2%; 3. GPT-5.6 Sol at 83%.

    Which mmmu-pro model is the cheapest?

    Llama 3.2 11B Instruct is the cheapest scored model on this mmmu-pro leaderboard at $0.05/M input and $0.05/M output ($0.10 blended). Gemini 3.5 Flash still leads MMMU-Pro at 83.6%.

    Should I always pick the #1 MMMU-Pro model?

    Not automatically. Gemini 3.5 Flash leads MMMU-Pro, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh MMMU-Pro against input/output price, context window, and related evals.

    How often is the MMMU-Pro leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is MMMU-Pro?

    MMMU-Pro — a harder subset of MMMU with more complex multimodal reasoning. This page ranks models that have published a MMMU-Pro score, with live API token prices on the same row.

    Where is the MMMU-Pro leaderboard?

    This page is the MMMU-Pro leaderboard. Models are sorted by MMMU-Pro, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this MMMU-Pro ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published MMMU-Pro score next to live API $/M so you can pick a production SKU, not only a trophy number.