GSM8K

    ๐Ÿ† Leaderboard

    As of August 20, 2026, MiMo-V2.5-Pro is #1 for GSM8K at 99.6%. Ranked by the GSM8K score 36 models in this index have a published GSM8K score. Methodology: GSM8K (https://github.com/openai/grade-school-math). GSM8K leaderboard: rank models by GSM8K next to live API token prices. Official methodology: GSM8K (https://github.com/openai/grade-school-math).

    Updated August 20, 2026282 models33 providers
    MiMo-V2.5-ProOSS
    Xiaomi ยท Open Source
    99.6%$1.30
    Kimi K2 InstructOSS
    Moonshot AI ยท Open Source
    97.3%$1.00
    o1
    OpenAI ยท Proprietary
    97.1%$75.00
    Llama 3.1 405B InstructOSS
    Meta ยท Open Source
    96.8%$1.78
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    96.4%$18.00
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    96.4%$18.00
    Gemma 3 27BOSS
    Google ยท Open Source
    95.9%$0.30
    Qwen2.5 72B InstructOSS
    Qwen ยท Open Source
    95.8%$0.75
    DeepSeek-V2.5OSS
    DeepSeek ยท Open Source
    95.1%$0.42
    Claude 3 Opus
    Anthropic ยท Proprietary
    95%$90.00
    Nova Pro
    Amazon ยท Proprietary
    94.8%$4.00
    Nova Lite
    Amazon ยท Proprietary
    94.5%$0.30
    Gemma 3 12BOSS
    Google ยท Open Source
    94.4%$0.15
    Qwen3 235B A22BOSS
    Qwen ยท Open Source
    94.4%$0.20
    Mistral Large 2OSS
    Mistral ยท Open Source ยท via Mistral AI
    93%$8.00
    Nova Micro
    Amazon ยท Proprietary
    92.3%$0.17
    Claude 3 Sonnet
    Anthropic ยท Proprietary
    92.3%$18.00
    Qwen2.5 7B InstructOSS
    Qwen ยท Open Source
    91.6%$0.60
    GPT-4o mini
    OpenAI ยท Proprietary
    91.3%$0.75
    Qwen2.5-Coder 32B InstructOSS
    Qwen ยท Open Source
    91.1%$0.18
    Gemini 1.5 Pro
    Google ยท Proprietary
    90.8%$12.50
    GPT-4
    OpenAI ยท Proprietary
    90.0%$90.00
    Gemma 3 4BOSS
    Google ยท Open Source
    89.2%$0.06
    Claude 3 Haiku
    Anthropic ยท Proprietary
    88.9%$1.50
    Phi 4 MiniOSS
    Microsoft ยท Open Source ยท via OpenRouter
    88.6%$0.43
    Jamba 1.5 LargeOSS
    AI21 Labs ยท Open Source
    87%$10.00
    Phi-3.5-mini-instructOSS
    Microsoft ยท Open Source
    86.2%$0.20
    Gemini 1.5 Flash
    Google ยท Proprietary
    86.2%$0.75
    Llama 3.1 8B InstructOSS
    Meta ยท Open Source
    82.4%$0.06
    Granite 3.3 8B InstructOSS
    IBM ยท Open Source
    80.9%$1.00
    Llama 3.2 3B InstructOSS
    Meta ยท Open Source
    77.7%$0.03lowest
    Jamba 1.5 MiniOSS
    AI21 Labs ยท Open Source
    75.8%$0.60
    Gemma 2 27BOSS
    Google ยท Open Source ยท via OpenRouter
    74%$1.30
    Command R+OSS
    Cohere ยท Open Source
    70.7%$1.25
    GPT-3.5 Turbo
    OpenAI ยท Proprietary
    57.8%$2.00
    ERNIE 4.5
    Baidu ยท Proprietary
    25.2%$4.40
    ChatGPT-4o Latest
    OpenAI ยท Proprietary
    โ€”$12.50
    Claude 3.5 Haiku
    Anthropic ยท Proprietary
    โ€”$4.80
    Claude 3.7 Sonnet
    Anthropic ยท Proprietary
    โ€”$18.00
    Claude Fable 5
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Haiku 4.5
    Anthropic ยท Proprietary
    โ€”$6.00
    Claude Mythos 5
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Mythos Preview
    Anthropic ยท Proprietary
    โ€”$60.00
    Claude Opus 4
    Anthropic ยท Proprietary
    โ€”$90.00
    Claude Opus 4.1
    Anthropic ยท Proprietary
    โ€”$90.00
    Claude Opus 4.5
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 4.6
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 4.7
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 4.8
    Anthropic ยท Proprietary
    โ€”$30.00
    Claude Opus 5
    Anthropic ยท Proprietary
    โ€”$30.00
    Showing 1โ€“50 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads GSM8K right now?

    As of August 20, 2026, MiMo-V2.5-Pro by Xiaomi is #1 for GSM8K at 99.6%. Ranked by the GSM8K score This board also tracks GSM8K. Next on the same board: Kimi K2 Instruct and o1. This gsm8k leaderboard ranks models by GSM8K. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: GSM8K (https://github.com/openai/grade-school-math); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for gsm8k. Ranked by the GSM8K score Input and output are dollars per million tokens.
    RankModelGSM8KInput /MOutput /M
    1MiMo-V2.5-Pro99.6%$0.43$0.87
    2Kimi K2 Instruct97.3%$0.50$0.50
    3o197.1%$15.00$60.00
    4Llama 3.1 405B Instruct96.8%$0.89$0.89
    5Claude 3.5 Sonnet96.4%$3.00$15.00
    6Claude 3.5 Sonnet96.4%$3.00$15.00
    7Gemma 3 27B95.9%$0.10$0.20
    8Qwen2.5 72B Instruct95.8%$0.35$0.40

    GSM8K FAQ

    Who ranks #1 on the GSM8K leaderboard?

    As of August 20, 2026, MiMo-V2.5-Pro by Xiaomi ranks #1 on GSM8K at 99.6%. API pricing is $0.43/M input and $0.87/M output.

    What are the top models on GSM8K?

    The current GSM8K ranking as of August 20, 2026 is 1. MiMo-V2.5-Pro at 99.6%; 2. Kimi K2 Instruct at 97.3%; 3. o1 at 97.1%.

    Which gsm8k model is the cheapest?

    Llama 3.2 3B Instruct is the cheapest scored model on this gsm8k leaderboard at $0.01/M input and $0.02/M output ($0.03 blended). MiMo-V2.5-Pro still leads GSM8K at 99.6%.

    Should I always pick the #1 GSM8K model?

    Not automatically. MiMo-V2.5-Pro leads GSM8K, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh GSM8K against input/output price, context window, and related evals.

    How often is the GSM8K leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is GSM8K?

    GSM8K is a public LLM eval (the GSM8K score). This page ranks models that have published a score, next to live API prices. Official methodology: GSM8K (https://github.com/openai/grade-school-math).

    Where is the GSM8K leaderboard?

    This page is the GSM8K leaderboard. Models are sorted by GSM8K, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this GSM8K ranking different from the official board?

    The official GSM8K page owns the methodology. This page keeps the published GSM8K score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://github.com/openai/grade-school-math