MATH 500

    🏆 Leaderboard

    As of August 20, 2026, LongCat-Flash-Thinking is #1 for MATH 500 at 99.2%. Ranked by MATH-500: textbook and contest math problems. 42 models in this index have a published MATH 500 score. MATH 500 leaderboard with live API prices. MATH-500 — a 500-problem subset of MATH covering competition-level mathematical reasoning.

    Updated August 20, 2026282 models33 providers
    LongCat-Flash-ThinkingOSS
    Meituan · Open Source
    99.2%$1.50
    GLM-4.5OSS
    Z AI · Open Source
    98.2%$2.80
    GLM-4.5-AirOSS
    Z AI · Open Source
    98.1%$1.30
    o3-mini
    OpenAI · Proprietary
    97.9%$5.50
    Nemotron Nano 9B v2OSS
    NVIDIA · Open Source · via OpenRouter
    97.8%
    Kimi K2-Instruct-0905OSS
    Moonshot AI · Open Source · via OpenRouter
    97.4%$3.10
    Kimi K2 InstructOSS
    Moonshot AI · Open Source
    97.4%$1.00
    DeepSeek-R1OSS
    DeepSeek · Open Source
    97.3%$2.74
    MiniMax M1 80KOSS
    MiniMax · Open Source
    96.8%$2.75
    LongCat-Flash-LiteOSS
    Meituan · Open Source
    96.8%$0.50
    LongCat-Flash-ChatOSS
    Meituan · Open Source
    96.4%$1.50
    Gemini 3 Pro
    Google · Proprietary
    96.4%$14.00
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    96.2%$18.00
    GPT-5
    OpenAI · Proprietary
    96%$11.25
    GPT-5 mini
    OpenAI · Proprietary
    94.8%$2.25
    GPT OSS 120BOSS
    OpenAI · Open Source
    94.8%$0.54
    Qwen3 235B A22BOSS
    Qwen · Open Source
    94.6%$0.20
    o3
    OpenAI · Proprietary
    94.6%$10.00
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek · Open Source
    94.5%$0.50
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek · Open Source
    94.3%$0.30
    o4-mini
    OpenAI · Proprietary
    94.2%$5.50
    GPT OSS 20BOSS
    OpenAI · Open Source
    94.2%$0.25
    DeepSeek-V3 0324OSS
    DeepSeek · Open Source
    94%$1.42
    GPT-5 nano
    OpenAI · Proprietary
    93.8%$0.45
    Claude Opus 4.1
    Anthropic · Proprietary
    93%$90.00
    QwQ-32B-PreviewOSS
    Qwen · Open Source
    90.6%$0.75
    o1
    OpenAI · Proprietary
    90.4%$75.00
    Claude Opus 4
    Anthropic · Proprietary
    90.4%$90.00
    Claude Sonnet 4
    Anthropic · Proprietary
    90.3%$18.00
    DeepSeek-V3OSS
    DeepSeek · Open Source
    90.2%$1.37
    o1-mini
    OpenAI · Proprietary
    90%$15.00
    Grok-3
    xAI · Proprietary
    89.8%$18.00
    Gemini 2.0 Flash
    Google · Proprietary
    89.7%$0.50
    MiniMax M2.1OSS
    MiniMax · Open Source
    89%$1.50
    GPT-4.1 mini
    OpenAI · Proprietary
    88%$2.00
    GPT-4.1
    OpenAI · Proprietary
    87.2%$10.00
    GPT-4.1 nano
    OpenAI · Proprietary
    80.2%$0.50
    GPT-4o
    OpenAI · Proprietary
    75.2%$12.50
    GPT-4o mini
    OpenAI · Proprietary
    72.6%$0.75
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    72.4%$18.00
    Granite 3.3 8B InstructOSS
    IBM · Open Source
    69.0%$1.00
    Claude 3.5 Haiku
    Anthropic · Proprietary
    64.2%$4.80
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Claude 3 Haiku
    Anthropic · Proprietary
    $1.50
    Claude 3 Opus
    Anthropic · Proprietary
    $90.00
    Claude 3 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude Fable 5
    Anthropic · Proprietary
    $60.00
    Claude Haiku 4.5
    Anthropic · Proprietary
    $6.00
    Claude Mythos 5
    Anthropic · Proprietary
    $60.00
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads MATH 500 right now?

    As of August 20, 2026, LongCat-Flash-Thinking by Meituan is #1 for MATH 500 at 99.2%. Ranked by MATH-500: textbook and contest math problems. This board also tracks MATH 500. Next on the same board: GLM-4.5 and GLM-4.5-Air. This math 500 leaderboard ranks models by MATH 500. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for math 500. Ranked by MATH-500: textbook and contest math problems. Input and output are dollars per million tokens.
    RankModelMATH 500Input /MOutput /M
    1LongCat-Flash-Thinking99.2%$0.30$1.20
    2GLM-4.598.2%$0.60$2.20
    3GLM-4.5-Air98.1%$0.20$1.10
    4o3-mini97.9%$1.10$4.40
    5Nemotron Nano 9B v297.8%
    6Kimi K2-Instruct-090597.4%$0.60$2.50
    7Kimi K2 Instruct97.4%$0.50$0.50
    8DeepSeek-R197.3%$0.55$2.19

    MATH 500 FAQ

    Who ranks #1 on the MATH 500 leaderboard?

    As of August 20, 2026, LongCat-Flash-Thinking by Meituan ranks #1 on MATH 500 at 99.2%. API pricing is $0.30/M input and $1.20/M output.

    What are the top models on MATH 500?

    The current MATH 500 ranking as of August 20, 2026 is 1. LongCat-Flash-Thinking at 99.2%; 2. GLM-4.5 at 98.2%; 3. GLM-4.5-Air at 98.1%.

    Which math 500 model is the cheapest?

    Qwen3 235B A22B is the cheapest scored model on this math 500 leaderboard at $0.10/M input and $0.10/M output ($0.20 blended). LongCat-Flash-Thinking still leads MATH 500 at 99.2%.

    Should I always pick the #1 MATH 500 model?

    Not automatically. LongCat-Flash-Thinking leads MATH 500, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh MATH 500 against input/output price, context window, and related evals.

    How often is the MATH 500 leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is MATH 500?

    MATH-500 — a 500-problem subset of MATH covering competition-level mathematical reasoning. This page ranks models that have published a MATH 500 score, with live API token prices on the same row.

    Where is the MATH 500 leaderboard?

    This page is the MATH 500 leaderboard. Models are sorted by MATH 500, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this MATH 500 ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published MATH 500 score next to live API $/M so you can pick a production SKU, not only a trophy number.