MMLU-Pro

    🏆 Leaderboard

    As of August 20, 2026, Claude Opus 5 is #1 for MMLU-Pro at 91.6%. Ranked by MMLU-Pro, a hard multiple-choice knowledge quiz. It is not a writing test. 165 models in this index have a published MMLU-Pro score. Methodology: MMLU-Pro (https://github.com/TIGER-AI-Lab/MMLU-Pro). MMLU-Pro leaderboard: rank models by MMLU-Pro next to live API token prices. Official methodology: MMLU-Pro (https://github.com/TIGER-AI-Lab/MMLU-Pro).

    Updated August 20, 2026282 models33 providers
    Claude Opus 5
    Anthropic · Proprietary
    91.6%$30.00
    Claude Fable 5
    Anthropic · Proprietary
    91.5%$60.00
    Gemini 3.1 Pro
    Google · Proprietary
    91.0%$17.50
    Sakana Namazu
    Sakana AI · Proprietary
    90.3%$4.95
    Gemini 3.7 Flash
    Google · Proprietary
    90.1%$4.50
    Gemini 3 Pro
    Google · Proprietary
    90.1%$14.00
    Claude Opus 4.7
    Anthropic · Proprietary
    89.9%$30.00
    Qwen3.7 Max
    Qwen · Proprietary
    89.6%$5.00
    Claude Opus 4.8
    Anthropic · Proprietary
    89.6%$30.00
    Gemini 3.5 Flash
    Google · Proprietary
    89.5%$10.50
    Grok 4.6
    xAI · Proprietary
    89.4%$8.00
    Gemini 3.6 Flash
    Google · Proprietary
    89.3%$4.50
    Grok 4.5
    xAI · Proprietary
    89.2%$8.00
    GPT-5.6 Sol
    OpenAI · Proprietary
    89.1%$35.00
    Muse Spark 1.1
    Meta · Proprietary
    88.7%$5.50
    Qwen3.8 MaxOSS
    Qwen · Open Source
    88.6%$8.00
    Gemini 3 Flash
    Google · Proprietary
    88.6%$3.50
    Qwen3.7-Plus
    Qwen · Proprietary
    88.5%$1.60
    Qwen3.6 Plus
    Qwen · Proprietary
    88.5%$3.50
    Muse Spark 1.2
    Meta · Proprietary
    88.3%$5.50
    GPT-5.5
    OpenAI · Proprietary
    88.1%$35.00
    MiniMax M2.1OSS
    MiniMax · Open Source
    88%$1.50
    Kimi K3OSS
    Moonshot AI · Open Source
    88.0%$18.00
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    87.8%$4.20
    Kimi K2.6OSS
    Moonshot AI · Open Source
    87.6%$4.93
    Claude Sonnet 5
    Anthropic · Proprietary
    87.5%$12.00
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    87.5%$5.22
    GPT-5.4
    OpenAI · Proprietary
    87.5%$17.50
    Claude Sonnet 4.6
    Anthropic · Proprietary
    87.3%$18.00
    DeepSeek-V4-Pro-0813OSS
    DeepSeek · Open Source
    87.3%$5.28
    Claude Opus 4.1
    Anthropic · Proprietary
    87.2%$90.00
    Kimi K2.5OSS
    Moonshot AI · Open Source
    87.1%$3.68
    GLM-5.1OSS
    Z AI · Open Source
    86.9%$5.80
    Nemotron 3 Ultra (550B A55B)OSS
    NVIDIA · Open Source · via OpenRouter
    86.8%$2.70
    GLM-5.2OSS
    Z AI · Open Source
    86.7%$5.80
    Qwen3.5-122B-A10BOSS
    Qwen · Open Source
    86.7%$3.60
    GPT-5.6 Terra
    OpenAI · Proprietary
    86.7%$14.00
    GPT-5
    OpenAI · Proprietary
    86.5%$11.25
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    86.4%$0.30
    GPT-5.1
    OpenAI · Proprietary
    86.4%$11.25
    Solar Pro 4
    Upstage · Proprietary
    86.3%$1.50
    Gemini 3.1 Flash-Lite
    Google · Proprietary
    86.2%$1.75
    GPT-5.2
    OpenAI · Proprietary
    86.2%$15.75
    DeepSeek-V4-Flash-0731OSS
    DeepSeek · Open Source
    86.2%$1.76
    Qwen3.6-27BOSS
    Qwen · Open Source
    86.2%$4.20
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    86.2%$0.42
    Claude Opus 4
    Anthropic · Proprietary
    86.2%$90.00
    Qwen3.5-27BOSS
    Qwen · Open Source
    86.1%$2.70
    GPT-5.6 Luna
    OpenAI · Proprietary
    86.0%$1.40
    Grok 4.3
    xAI · Proprietary
    85.8%$3.75
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads MMLU-Pro right now?

    As of August 20, 2026, Claude Opus 5 by Anthropic is #1 for MMLU-Pro at 91.6%. Ranked by MMLU-Pro, a hard multiple-choice knowledge quiz. It is not a writing test. This board also tracks MMLU-Pro. Next on the same board: Claude Fable 5 and Gemini 3.1 Pro. This mmlu-pro leaderboard ranks models by MMLU-Pro. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: MMLU-Pro (https://github.com/TIGER-AI-Lab/MMLU-Pro); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for mmlu-pro. Ranked by MMLU-Pro, a hard multiple-choice knowledge quiz. It is not a writing test. Input and output are dollars per million tokens.
    RankModelMMLU-ProInput /MOutput /M
    1Claude Opus 591.6%$5.00$25.00
    2Claude Fable 591.5%$10.00$50.00
    3Gemini 3.1 Pro91.0%$2.50$15.00
    4Sakana Namazu90.3%$0.95$4.00
    5Gemini 3.7 Flash90.1%$0.75$3.75
    6Gemini 3 Pro90.1%$2.00$12.00
    7Claude Opus 4.789.9%$5.00$25.00
    8Qwen3.7 Max89.6%$1.25$3.75

    MMLU-Pro FAQ

    Who ranks #1 on the MMLU-Pro leaderboard?

    As of August 20, 2026, Claude Opus 5 by Anthropic ranks #1 on MMLU-Pro at 91.6%. API pricing is $5.00/M input and $25.00/M output.

    What are the top models on MMLU-Pro?

    The current MMLU-Pro ranking as of August 20, 2026 is 1. Claude Opus 5 at 91.6%; 2. Claude Fable 5 at 91.5%; 3. Gemini 3.1 Pro at 91.0%.

    Which mmlu-pro model is the cheapest?

    Llama 3.1 8B Instruct is the cheapest scored model on this mmlu-pro leaderboard at $0.03/M input and $0.03/M output ($0.06 blended). Claude Opus 5 still leads MMLU-Pro at 91.6%.

    Should I always pick the #1 MMLU-Pro model?

    Not automatically. Claude Opus 5 leads MMLU-Pro, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh MMLU-Pro against input/output price, context window, and related evals.

    How often is the MMLU-Pro leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is MMLU-Pro?

    MMLU-Pro is a public LLM eval (MMLU-Pro, a hard multiple-choice knowledge quiz. It is not a writing test.). This page ranks models that have published a score, next to live API prices. Official methodology: MMLU-Pro (https://github.com/TIGER-AI-Lab/MMLU-Pro).

    Where is the MMLU-Pro leaderboard?

    This page is the MMLU-Pro leaderboard. Models are sorted by MMLU-Pro, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this MMLU-Pro ranking different from the official board?

    The official MMLU-Pro page owns the methodology. This page keeps the published MMLU-Pro score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://github.com/TIGER-AI-Lab/MMLU-Pro