MMLU-Redux

    ๐Ÿ† Leaderboard

    As of August 20, 2026, Qwen3.7 Max is #1 for MMLU-Redux at 95%. Ranked by the MMLU-Redux score 37 models in this index have a published MMLU-Redux score. MMLU-Redux leaderboard: rank models by MMLU-Redux next to live API token prices.

    Updated August 20, 2026282 models33 providers
    Qwen3.7 Max
    Qwen ยท Proprietary
    95%$5.00
    Qwen3.5-397B-A17BOSS
    Qwen ยท Open Source
    94.9%$4.20
    Qwen3.7-Plus
    Qwen ยท Proprietary
    94.5%$1.60
    Qwen3.6 Plus
    Qwen ยท Proprietary
    94.5%$3.50
    Kimi K2-Thinking-0905OSS
    Moonshot AI ยท Open Source
    94.4%$2.47
    Qwen3.5-122B-A10BOSS
    Qwen ยท Open Source
    94%$3.60
    Qwen3-235B-A22B-Thinking-2507OSS
    Qwen ยท Open Source
    93.8%$3.30
    Qwen3 VL 235B A22B ThinkingOSS
    Qwen ยท Open Source
    93.7%$3.94
    Qwen3.6-27BOSS
    Qwen ยท Open Source
    93.5%$4.20
    DeepSeek-R1-0528OSS
    DeepSeek ยท Open Source
    93.4%$2.74
    Qwen3.6-35B-A3BOSS
    Qwen ยท Open Source ยท via OpenRouter
    93.3%$1.14
    Qwen3.5-35B-A3BOSS
    Qwen ยท Open Source
    93.3%$2.25
    Qwen3.5-27BOSS
    Qwen ยท Open Source
    93.2%$2.70
    Qwen3-235B-A22B-Instruct-2507OSS
    Qwen ยท Open Source
    93.1%$0.95
    MiMo-V2.5-ProOSS
    Xiaomi ยท Open Source
    92.8%$1.30
    Kimi K2-Instruct-0905OSS
    Moonshot AI ยท Open Source ยท via OpenRouter
    92.7%$3.10
    Kimi K2 InstructOSS
    Moonshot AI ยท Open Source
    92.7%$1.00
    Qwen3-Next-80B-A3B-ThinkingOSS
    Qwen ยท Open Source
    92.5%$1.65
    Qwen3 VL 235B A22B InstructOSS
    Qwen ยท Open Source
    92.2%$1.79
    DeepSeek-V3.1OSS
    DeepSeek ยท Open Source
    91.8%$1.27
    Qwen3.5-9BOSS
    Qwen ยท Open Source ยท via OpenRouter
    91.1%$0.25
    Qwen3-Next-80B-A3B-InstructOSS
    Qwen ยท Open Source
    90.9%$1.65
    Qwen3 VL 30B A3B ThinkingOSS
    Qwen ยท Open Source
    90.9%$1.20
    Qwen3 VL 32B InstructOSS
    Qwen ยท Open Source ยท via OpenRouter
    89.8%$0.52
    LongCat-Flash-ThinkingOSS
    Meituan ยท Open Source
    89.3%$1.50
    DeepSeek-V3OSS
    DeepSeek ยท Open Source
    89.1%$1.37
    Qwen3 VL 8B ThinkingOSS
    Qwen ยท Open Source
    88.8%$2.27
    Qwen3 VL 30B A3B InstructOSS
    Qwen ยท Open Source
    88.4%$0.90
    Qwen3 235B A22BOSS
    Qwen ยท Open Source
    87.4%$0.20
    Qwen2.5 72B InstructOSS
    Qwen ยท Open Source
    86.8%$0.75
    Qwen3 VL 4B ThinkingOSS
    Qwen ยท Open Source
    86%$1.10
    Qwen3 VL 8B InstructOSS
    Qwen ยท Open Source
    84.9%$0.58
    Mistral Large 3OSS
    Mistral ยท Open Source ยท via Mistral AI
    82%$7.00
    Qwen3 VL 4B InstructOSS
    Qwen ยท Open Source
    81.5%$0.70
    Qwen2.5-Coder 32B InstructOSS
    Qwen ยท Open Source
    77.5%$0.18
    Qwen2.5 7B InstructOSS
    Qwen ยท Open Source
    75.4%$0.60
    ERNIE 4.5
    Baidu ยท Proprietary
    43.2%$4.40
    ChatGPT-4o Latest
    OpenAI ยท Proprietary
    โ€”$12.50
    Claude 3 Haiku
    Anthropic ยท Proprietary
    โ€”$1.50
    Claude 3 Opus
    Anthropic ยท Proprietary
    โ€”$90.00
    Claude 3 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude 3.5 Haiku
    Anthropic ยท Proprietary
    โ€”$4.80
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude 3.7 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude Fable 5
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Haiku 4.5
    Anthropic ยท Proprietary
    โ€”$6.00
    Claude Mythos 5
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Mythos Preview
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Opus 4
    Anthropic ยท Proprietary
    โ€”$90.00
    Showing 1โ€“50 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads MMLU-Redux right now?

    As of August 20, 2026, Qwen3.7 Max by Qwen is #1 for MMLU-Redux at 95%. Ranked by the MMLU-Redux score This board also tracks MMLU-Redux. Next on the same board: Qwen3.5-397B-A17B and Qwen3.7-Plus. This mmlu-redux leaderboard ranks models by MMLU-Redux. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for mmlu-redux. Ranked by the MMLU-Redux score Input and output are dollars per million tokens.
    RankModelMMLU-ReduxInput /MOutput /M
    1Qwen3.7 Max95%$1.25$3.75
    2Qwen3.5-397B-A17B94.9%$0.60$3.60
    3Qwen3.7-Plus94.5%$0.32$1.28
    4Qwen3.6 Plus94.5%$0.50$3.00
    5Kimi K2-Thinking-090594.4%$0.47$2.00
    6Qwen3.5-122B-A10B94%$0.40$3.20
    7Qwen3-235B-A22B-Thinking-250793.8%$0.30$3.00
    8Qwen3 VL 235B A22B Thinking93.7%$0.45$3.49

    MMLU-Redux FAQ

    Who ranks #1 on the MMLU-Redux leaderboard?

    As of August 20, 2026, Qwen3.7 Max by Qwen ranks #1 on MMLU-Redux at 95%. API pricing is $1.25/M input and $3.75/M output.

    What are the top models on MMLU-Redux?

    The current MMLU-Redux ranking as of August 20, 2026 is 1. Qwen3.7 Max at 95%; 2. Qwen3.5-397B-A17B at 94.9%; 3. Qwen3.7-Plus at 94.5%.

    Which mmlu-redux model is the cheapest?

    Qwen2.5-Coder 32B Instruct is the cheapest scored model on this mmlu-redux leaderboard at $0.09/M input and $0.09/M output ($0.18 blended). Qwen3.7 Max still leads MMLU-Redux at 95%.

    Should I always pick the #1 MMLU-Redux model?

    Not automatically. Qwen3.7 Max leads MMLU-Redux, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh MMLU-Redux against input/output price, context window, and related evals.

    How often is the MMLU-Redux leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is MMLU-Redux?

    MMLU-Redux is a public LLM eval (the MMLU-Redux score). This page ranks models that have published a score, next to live API prices.

    Where is the MMLU-Redux leaderboard?

    This page is the MMLU-Redux leaderboard. Models are sorted by MMLU-Redux, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this MMLU-Redux ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published MMLU-Redux score next to live API $/M so you can pick a production SKU, not only a trophy number.