MMMLU

    ๐Ÿ† Leaderboard

    As of August 20, 2026, Claude Mythos Preview is #1 for MMMLU at 92.7%. Ranked by the MMMLU score 38 models in this index have a published MMMLU score. MMMLU leaderboard with live API prices. Multilingual MMLU โ€” MMLU translated into 14+ languages to test cross-lingual knowledge.

    Updated August 20, 2026282 models33 providers
    Claude Mythos Preview
    Anthropic ยท Proprietary
    92.7%$60.00
    Gemini 3.1 Pro
    Google ยท Proprietary
    92.6%$17.50
    Gemini 3 Pro
    Google ยท Proprietary
    91.8%$14.00
    Gemini 3 Flash
    Google ยท Proprietary
    91.8%$3.50
    Claude Opus 4.7
    Anthropic ยท Proprietary
    91.5%$30.00
    Claude Opus 4.6
    Anthropic ยท Proprietary
    91.1%$30.00
    Claude Opus 4.5
    Anthropic ยท Proprietary
    90.8%$30.00
    Qwen3.7 Max
    Qwen ยท Proprietary
    90.3%$5.00
    GPT-5.2
    OpenAI ยท Proprietary
    89.6%$15.75
    Qwen3.6 Plus
    Qwen ยท Proprietary
    89.5%$3.50
    Claude Opus 4.1
    Anthropic ยท Proprietary
    89.5%$90.00
    Claude Sonnet 4.6
    Anthropic ยท Proprietary
    89.3%$18.00
    Gemini 2.5 Pro
    Google ยท Proprietary
    89.2%$11.25
    Claude Sonnet 4.5
    Anthropic ยท Proprietary
    89.1%$18.00
    Qwen3.7-Plus
    Qwen ยท Proprietary
    89%$1.60
    Gemini 3.1 Flash-Lite
    Google ยท Proprietary
    88.9%$1.75
    Claude Opus 4
    Anthropic ยท Proprietary
    88.8%$90.00
    Qwen3.5-397B-A17BOSS
    Qwen ยท Open Source
    88.5%$4.20
    Gemma 4 31BOSS
    Google ยท Open Source
    88.4%$0.54
    o1
    OpenAI ยท Proprietary
    87.7%$75.00
    GPT-4.1
    OpenAI ยท Proprietary
    87.3%$10.00
    Qwen3.5-122B-A10BOSS
    Qwen ยท Open Source
    86.7%$3.60
    Qwen3 235B A22BOSS
    Qwen ยท Open Source
    86.7%$0.20
    Claude Sonnet 4
    Anthropic ยท Proprietary
    86.5%$18.00
    Gemma 4 26B-A4BOSS
    Google ยท Open Source
    86.3%$0.53
    Claude 3.7 Sonnet
    Anthropic ยท Proprietary
    86.1%$18.00
    Qwen3.5-27BOSS
    Qwen ยท Open Source
    85.9%$2.70
    K-EXAONE-236B-A23B
    LG AI Research ยท Proprietary
    85.7%$1.60
    Mistral Large 3 (675B Instruct 2512)OSS
    Mistral ยท Open Source ยท via Mistral AI
    85.5%$2.00
    Qwen3.5-35B-A3BOSS
    Qwen ยท Open Source
    85.2%$2.25
    GPT OSS 120B HighOSS
    OpenAI ยท Open Source
    83.8%$0.60
    Claude Haiku 4.5
    Anthropic ยท Proprietary
    83%$6.00
    GPT-4o
    OpenAI ยท Proprietary
    81.4%$12.50
    Qwen3.5-9BOSS
    Qwen ยท Open Source ยท via OpenRouter
    81.2%$0.25
    GPT-4.1 mini
    OpenAI ยท Proprietary
    78.5%$2.00
    Mistral Large 3OSS
    Mistral ยท Open Source ยท via Mistral AI
    74.2%$7.00
    GPT-4.1 nano
    OpenAI ยท Proprietary
    66.9%$0.50
    Phi-3.5-mini-instructOSS
    Microsoft ยท Open Source
    55.4%$0.20
    ChatGPT-4o Latest
    OpenAI ยท Proprietary
    โ€”$12.50
    Claude 3 Haiku
    Anthropic ยท Proprietary
    โ€”$1.50
    Claude 3 Opus
    Anthropic ยท Proprietary
    โ€”$90.00
    Claude 3 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude 3.5 Haiku
    Anthropic ยท Proprietary
    โ€”$4.80
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude Fable 5
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Mythos 5
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Opus 4.8
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 5
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Sonnet 5
    Anthropic ยท Proprietary
    โ€”$12.00
    Showing 1โ€“50 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads MMMLU right now?

    As of August 20, 2026, Claude Mythos Preview by Anthropic is #1 for MMMLU at 92.7%. Ranked by the MMMLU score This board also tracks MMMLU. Next on the same board: Gemini 3.1 Pro and Gemini 3 Flash. This mmmlu leaderboard ranks models by MMMLU. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for mmmlu. Ranked by the MMMLU score Input and output are dollars per million tokens.
    RankModelMMMLUInput /MOutput /M
    1Claude Mythos Preview92.7%$10.00$50.00
    2Gemini 3.1 Pro92.6%$2.50$15.00
    3Gemini 3 Flash91.8%$0.50$3.00
    4Gemini 3 Pro91.8%$2.00$12.00
    5Claude Opus 4.791.5%$5.00$25.00
    6Claude Opus 4.691.1%$5.00$25.00
    7Claude Opus 4.590.8%$5.00$25.00
    8Qwen3.7 Max90.3%$1.25$3.75

    MMMLU FAQ

    Who ranks #1 on the MMMLU leaderboard?

    As of August 20, 2026, Claude Mythos Preview by Anthropic ranks #1 on MMMLU at 92.7%. API pricing is $10.00/M input and $50.00/M output.

    What are the top models on MMMLU?

    The current MMMLU ranking as of August 20, 2026 is 1. Claude Mythos Preview at 92.7%; 2. Gemini 3.1 Pro at 92.6%; 3. Gemini 3 Flash at 91.8%.

    Which mmmlu model is the cheapest?

    Qwen3 235B A22B is the cheapest scored model on this mmmlu leaderboard at $0.10/M input and $0.10/M output ($0.20 blended). Claude Mythos Preview still leads MMMLU at 92.7%.

    Should I always pick the #1 MMMLU model?

    Not automatically. Claude Mythos Preview leads MMMLU, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh MMMLU against input/output price, context window, and related evals.

    How often is the MMMLU leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is MMMLU?

    Multilingual MMLU โ€” MMLU translated into 14+ languages to test cross-lingual knowledge. This page ranks models that have published a MMMLU score, with live API token prices on the same row.

    Where is the MMMLU leaderboard?

    This page is the MMMLU leaderboard. Models are sorted by MMMLU, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this MMMLU ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published MMMLU score next to live API $/M so you can pick a production SKU, not only a trophy number.