SWE-bench Pro (Public)

    🏆 Leaderboard

    As of August 20, 2026, GPT-5.4 is #1 for SWE-bench Pro (Public) at 88.4%. Ranked by the SWE-bench Pro (Public) score 10 models in this index have a published SWE-bench Pro (Public) score. Methodology: SWE-bench (https://www.swebench.com/). SWE-bench Pro (Public) leaderboard with live API prices. SWE-bench Pro (Public) — a harder curated subset of SWE-bench with more complex real-world engineering tasks. Official methodology: SWE-bench (https://www.swebench.com/).

    Updated August 20, 2026282 models33 providers
    GPT-5.4
    OpenAI · Proprietary
    88.4%$17.50
    GPT-5.3 Codex
    OpenAI · Proprietary
    84.8%$15.75
    Gemini 3 Pro
    Google · Proprietary
    84.8%$14.00
    Claude Mythos 5
    Anthropic · Proprietary
    80.3%$60.00
    Claude Fable 5
    Anthropic · Proprietary
    80%$60.00
    Gemini 3.6 Flash
    Google · Proprietary
    58.7%$4.50
    GPT-5.5
    OpenAI · Proprietary
    58.6%$35.00
    GPT-5.4 mini
    OpenAI · Proprietary
    54.4%$5.25
    Gemini 3.5 Flash-Lite
    Google · Proprietary
    54.2%$2.80
    GPT-5.4 nano
    OpenAI · Proprietary
    52.4%$1.45
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Claude 3 Haiku
    Anthropic · Proprietary
    $1.50
    Claude 3 Opus
    Anthropic · Proprietary
    $90.00
    Claude 3 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Haiku
    Anthropic · Proprietary
    $4.80
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude Haiku 4.5
    Anthropic · Proprietary
    $6.00
    Claude Mythos Preview
    Anthropic · Proprietary
    $60.00
    Claude Opus 4
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.1
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.5
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.6
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.7
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.8
    Anthropic · Proprietary
    $30.00
    Claude Opus 5
    Anthropic · Proprietary
    $30.00
    Claude Sonnet 4
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 4.5
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 4.6
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 5
    Anthropic · Proprietary
    $12.00
    Command A+OSS
    Cohere · Open Source
    $12.50
    Command R+OSS
    Cohere · Open Source
    $1.25
    Composer 2
    Cursor · Proprietary
    $3.00
    Composer 2 Fast
    Cursor · Proprietary
    $9.00
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek · Open Source
    $0.50
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek · Open Source
    $0.30
    DeepSeek-R1OSS
    DeepSeek · Open Source
    $2.74
    DeepSeek-R1-0528OSS
    DeepSeek · Open Source
    $2.74
    DeepSeek-V2.5OSS
    DeepSeek · Open Source
    $0.42
    DeepSeek-V3OSS
    DeepSeek · Open Source
    $1.37
    DeepSeek-V3 0324OSS
    DeepSeek · Open Source
    $1.42
    DeepSeek-V3.1OSS
    DeepSeek · Open Source
    $1.27
    DeepSeek-V3.2OSS
    DeepSeek · Open Source · via OpenRouter
    $0.57
    DeepSeek-V3.2 (Non-thinking)OSS
    DeepSeek · Open Source
    $0.70
    DeepSeek-V3.2-ExpOSS
    DeepSeek · Open Source
    $0.68
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    $0.30
    DeepSeek-V4-Flash-0731OSS
    DeepSeek · Open Source
    $1.76
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    $0.42
    DeepSeek-V4-Pro-0813OSS
    DeepSeek · Open Source
    $5.28
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads SWE-bench Pro (Public) right now?

    As of August 20, 2026, GPT-5.4 by OpenAI is #1 for SWE-bench Pro (Public) at 88.4%. Ranked by the SWE-bench Pro (Public) score This board also tracks SWE-bench Pro (Public). Next on the same board: GPT-5.3 Codex and Gemini 3 Pro. This swe-bench pro (public) leaderboard ranks models by SWE-bench Pro (Public). Scores come from public evals. Prices are the live API rates in the table above.

    Sources: SWE-bench (https://www.swebench.com/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for swe-bench pro (public). Ranked by the SWE-bench Pro (Public) score Input and output are dollars per million tokens.
    RankModelSWE-bench Pro (Public)Input /MOutput /M
    1GPT-5.488.4%$2.50$15.00
    2GPT-5.3 Codex84.8%$1.75$14.00
    3Gemini 3 Pro84.8%$2.00$12.00
    4Claude Mythos 580.3%$10.00$50.00
    5Claude Fable 580%$10.00$50.00
    6Gemini 3.6 Flash58.7%$0.75$3.75
    7GPT-5.558.6%$5.00$30.00
    8GPT-5.4 mini54.4%$0.75$4.50

    SWE-bench Pro (Public) FAQ

    Who ranks #1 on the SWE-bench Pro (Public) leaderboard?

    As of August 20, 2026, GPT-5.4 by OpenAI ranks #1 on SWE-bench Pro (Public) at 88.4%. API pricing is $2.50/M input and $15.00/M output.

    What are the top models on SWE-bench Pro (Public)?

    The current SWE-bench Pro (Public) ranking as of August 20, 2026 is 1. GPT-5.4 at 88.4%; 2. GPT-5.3 Codex at 84.8%; 3. Gemini 3 Pro at 84.8%.

    Which swe-bench pro (public) model is the cheapest?

    GPT-5.4 nano is the cheapest scored model on this swe-bench pro (public) leaderboard at $0.20/M input and $1.25/M output ($1.45 blended). GPT-5.4 still leads SWE-bench Pro (Public) at 88.4%.

    Should I always pick the #1 SWE-bench Pro (Public) model?

    Not automatically. GPT-5.4 leads SWE-bench Pro (Public), but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh SWE-bench Pro (Public) against input/output price, context window, and related evals.

    How often is the SWE-bench Pro (Public) leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is SWE-bench Pro (Public)?

    SWE-bench Pro (Public) — a harder curated subset of SWE-bench with more complex real-world engineering tasks. This page ranks models that have published a SWE-bench Pro (Public) score, with live API token prices on the same row. Official methodology: SWE-bench (https://www.swebench.com/).

    Where is the SWE-bench Pro (Public) leaderboard?

    This page is the SWE-bench Pro (Public) leaderboard. Models are sorted by SWE-bench Pro (Public), with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this SWE-bench Pro (Public) ranking different from the official board?

    The official SWE-bench page owns the methodology. This page keeps the published SWE-bench Pro (Public) score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://www.swebench.com/