Arena-Hard

    🏆 Leaderboard

    As of August 20, 2026, Qwen3 235B A22B is #1 for Arena-Hard at 95.6%. Ranked by the Arena-Hard score 20 models in this index have a published Arena-Hard score. Arena-Hard leaderboard: rank models by Arena-Hard next to live API token prices.

    Updated August 20, 2026282 models33 providers
    Qwen3 235B A22BOSS
    Qwen · Open Source
    95.6%$0.20
    Qwen3 32BOSS
    Qwen · Open Source
    93.8%$0.54
    Qwen3 30B A3BOSS
    Qwen · Open Source
    91%$0.54
    Mistral Small 3 24B InstructOSS
    Mistral · Open Source · via Mistral AI
    87.6%$0.21
    Qwen2.5 72B InstructOSS
    Qwen · Open Source
    81.2%$0.75
    DeepSeek-V2.5OSS
    DeepSeek · Open Source
    76.2%$0.42
    Phi 4OSS
    Microsoft · Open Source
    75.4%$0.21
    Ministral 8B InstructOSS
    Mistral · Open Source · via Mistral AI
    70.9%$0.20
    Jamba 1.5 LargeOSS
    AI21 Labs · Open Source
    65.4%$10.00
    Mistral Small 4OSS
    Mistral · Open Source · via Mistral AI
    58.3%$0.75
    Granite 3.3 8B InstructOSS
    IBM · Open Source
    57.6%$1.00
    Mistral Large 3OSS
    Mistral · Open Source · via Mistral AI
    55.1%$7.00
    MiniStral 3 (14B Instruct 2512)OSS
    Mistral · Open Source · via OpenRouter
    55.1%$0.40
    Qwen2.5 7B InstructOSS
    Qwen · Open Source
    52%$0.60
    Ministral 3 (8B Instruct 2512)OSS
    Mistral · Open Source · via OpenRouter
    50.9%$0.30
    Jamba 1.5 MiniOSS
    AI21 Labs · Open Source
    46.1%$0.60
    Mistral Small 3.2 24B InstructOSS
    Mistral · Open Source · via OpenRouter
    43.1%$0.28
    Phi-3.5-mini-instructOSS
    Microsoft · Open Source
    37%$0.20
    Phi 4 MiniOSS
    Microsoft · Open Source · via OpenRouter
    32.8%$0.43
    Ministral 3 (3B Instruct 2512)OSS
    Mistral · Open Source · via OpenRouter
    30.5%$0.20
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Claude 3 Haiku
    Anthropic · Proprietary
    $1.50
    Claude 3 Opus
    Anthropic · Proprietary
    $90.00
    Claude 3 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Haiku
    Anthropic · Proprietary
    $4.80
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude Fable 5
    Anthropic · Proprietary
    $60.00
    Claude Haiku 4.5
    Anthropic · Proprietary
    $6.00
    Claude Mythos 5
    Anthropic · Proprietary
    $60.00
    Claude Mythos Preview
    Anthropic · Proprietary
    $60.00
    Claude Opus 4
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.1
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.5
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.6
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.7
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.8
    Anthropic · Proprietary
    $30.00
    Claude Opus 5
    Anthropic · Proprietary
    $30.00
    Claude Sonnet 4
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 4.5
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 4.6
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 5
    Anthropic · Proprietary
    $12.00
    Command A+OSS
    Cohere · Open Source
    $12.50
    Command R+OSS
    Cohere · Open Source
    $1.25
    Composer 2
    Cursor · Proprietary
    $3.00
    Composer 2 Fast
    Cursor · Proprietary
    $9.00
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek · Open Source
    $0.50
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek · Open Source
    $0.30
    DeepSeek-R1OSS
    DeepSeek · Open Source
    $2.74
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads Arena-Hard right now?

    As of August 20, 2026, Qwen3 235B A22B by Qwen is #1 for Arena-Hard at 95.6%. Ranked by the Arena-Hard score This board also tracks Arena-Hard. Next on the same board: Qwen3 32B and Qwen3 30B A3B. This arena-hard leaderboard ranks models by Arena-Hard. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for arena-hard. Ranked by the Arena-Hard score Input and output are dollars per million tokens.
    RankModelArena-HardInput /MOutput /M
    1Qwen3 235B A22B95.6%$0.10$0.10
    2Qwen3 32B93.8%$0.10$0.44
    3Qwen3 30B A3B91%$0.10$0.44
    4Mistral Small 3 24B Instruct87.6%$0.07$0.14
    5Qwen2.5 72B Instruct81.2%$0.35$0.40
    6DeepSeek-V2.576.2%$0.14$0.28
    7Phi 475.4%$0.07$0.14
    8Ministral 8B Instruct70.9%$0.10$0.10

    Arena-Hard FAQ

    Who ranks #1 on the Arena-Hard leaderboard?

    As of August 20, 2026, Qwen3 235B A22B by Qwen ranks #1 on Arena-Hard at 95.6%. API pricing is $0.10/M input and $0.10/M output.

    What are the top models on Arena-Hard?

    The current Arena-Hard ranking as of August 20, 2026 is 1. Qwen3 235B A22B at 95.6%; 2. Qwen3 32B at 93.8%; 3. Qwen3 30B A3B at 91%.

    Which arena-hard model is the cheapest?

    Qwen3 235B A22B currently leads Arena-Hard and is also the cheapest scored model on this page, at $0.10/M input and $0.10/M output ($0.20 blended).

    Should I always pick the #1 Arena-Hard model?

    Not automatically. Qwen3 235B A22B leads Arena-Hard, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh Arena-Hard against input/output price, context window, and related evals.

    How often is the Arena-Hard leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is Arena-Hard?

    Arena-Hard is a public LLM eval (the Arena-Hard score). This page ranks models that have published a score, next to live API prices.

    Where is the Arena-Hard leaderboard?

    This page is the Arena-Hard leaderboard. Models are sorted by Arena-Hard, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this Arena-Hard ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published Arena-Hard score next to live API $/M so you can pick a production SKU, not only a trophy number.