SWE-bench Verified

    🏆 Leaderboard

    As of August 20, 2026, DeepSeek-V4-Pro-0813 is #1 for SWE-bench Verified at 96.4%. Ranked by the SWE-bench Verified score 137 models in this index have a published SWE-bench Verified score. Methodology: SWE-bench Verified (https://www.swebench.com/verified.html). SWE-bench Verified leaderboard with live API prices. SWE-bench Verified — a human-validated subset of SWE-bench ensuring each task is solvable and well-specified. Official methodology: SWE-bench Verified (https://www.swebench.com/verified.html).

    Updated August 20, 2026282 models33 providers
    DeepSeek-V4-Pro-0813OSS
    DeepSeek · Open Source
    96.4%$5.28
    GPT-5.6 Sol
    OpenAI · Proprietary
    96.2%$35.00
    Claude Opus 5
    Anthropic · Proprietary
    96%$30.00
    Grok 4.6
    xAI · Proprietary
    95.6%$8.00
    Claude Mythos 5
    Anthropic · Proprietary
    95.5%$60.00
    Claude Fable 5
    Anthropic · Proprietary
    95%$60.00
    Claude Mythos Preview
    Anthropic · Proprietary
    93.9%$60.00
    Kimi K3OSS
    Moonshot AI · Open Source
    93.4%$18.00
    GPT-5.6 Luna
    OpenAI · Proprietary
    93%$1.40
    DeepSeek-V4-Flash-0731OSS
    DeepSeek · Open Source
    88.8%$1.76
    Claude Opus 4.8
    Anthropic · Proprietary
    88.6%$30.00
    Claude Opus 4.7
    Anthropic · Proprietary
    87.6%$30.00
    Muse Spark 1.2
    Meta · Proprietary
    86.6%$5.50
    Grok 4.5
    xAI · Proprietary
    86.6%$8.00
    Qwen3.8 MaxOSS
    Qwen · Open Source
    85.6%$8.00
    Claude Sonnet 5
    Anthropic · Proprietary
    85.2%$12.00
    GPT-5.5
    OpenAI · Proprietary
    82.6%$35.00
    Muse Spark 1.1
    Meta · Proprietary
    82%$5.50
    Claude Opus 4.5
    Anthropic · Proprietary
    80.9%$30.00
    Gemini 3.7 Flash
    Google · Proprietary
    80.8%$4.50
    Claude Opus 4.6
    Anthropic · Proprietary
    80.8%$30.00
    Gemini 3.1 Pro
    Google · Proprietary
    80.6%$17.50
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    80.6%$5.22
    MiniMax M3OSS
    MiniMax · Open Source
    80.5%$1.50
    Qwen3.7 Max
    Qwen · Proprietary
    80.4%$5.00
    MiniMax M2.5OSS
    MiniMax · Open Source
    80.2%$1.50
    Kimi K2.6OSS
    Moonshot AI · Open Source
    80.2%$4.93
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    80.2%$1.50
    GPT-5.2
    OpenAI · Proprietary
    80%$15.75
    Gemini 3.6 Flash
    Google · Proprietary
    79.6%$4.50
    Claude Sonnet 4.6
    Anthropic · Proprietary
    79.6%$18.00
    Gemini 3.5 Flash
    Google · Proprietary
    79.3%$10.50
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    79%$0.42
    MiMo-V2.5-ProOSS
    Xiaomi · Open Source
    78.9%$1.30
    Qwen3.6 Plus
    Qwen · Proprietary
    78.8%$3.50
    GLM-5.2OSS
    Z AI · Open Source
    78.7%$5.80
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    78.6%$0.30
    Kimi K2.7 CodeOSS
    Moonshot AI · Open Source
    78.2%$4.93
    Hy3OSS
    Tencent · Open Source
    78%$0.66
    Gemini 3 Flash
    Google · Proprietary
    78%$3.50
    GLM-5OSS
    Z AI · Open Source
    77.8%$4.20
    Qwen3.7-Plus
    Qwen · Proprietary
    77.7%$1.60
    Mistral Medium 3.5OSS
    Mistral · Open Source · via Mistral AI
    77.6%$9.00
    Qwen3.6-27BOSS
    Qwen · Open Source
    77.2%$4.20
    GPT-5.4
    OpenAI · Proprietary
    76.9%$17.50
    Kimi K2.5OSS
    Moonshot AI · Open Source
    76.8%$3.68
    Seed 2.0 Pro
    ByteDance · Proprietary
    76.5%$3.50
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    76.4%$4.20
    GPT-5.1 Thinking
    OpenAI · Proprietary
    76.3%$11.25
    GPT-5.1 Instant
    OpenAI · Proprietary
    76.3%$11.25
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads SWE-bench Verified right now?

    As of August 20, 2026, DeepSeek-V4-Pro-0813 by DeepSeek is #1 for SWE-bench Verified at 96.4%. Ranked by the SWE-bench Verified score This board also tracks SWE-bench Verified. Next on the same board: GPT-5.6 Sol and Claude Opus 5. This swe-bench verified leaderboard ranks models by SWE-bench Verified. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: SWE-bench Verified (https://www.swebench.com/verified.html); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for swe-bench verified. Ranked by the SWE-bench Verified score Input and output are dollars per million tokens.
    RankModelSWE-bench VerifiedInput /MOutput /M
    1DeepSeek-V4-Pro-081396.4%$1.32$3.96
    2GPT-5.6 Sol96.2%$5.00$30.00
    3Claude Opus 596%$5.00$25.00
    4Grok 4.695.6%$2.00$6.00
    5Claude Mythos 595.5%$10.00$50.00
    6Claude Fable 595%$10.00$50.00
    7Claude Mythos Preview93.9%$10.00$50.00
    8Kimi K393.4%$3.00$15.00

    SWE-bench Verified FAQ

    Who ranks #1 on the SWE-bench Verified leaderboard?

    As of August 20, 2026, DeepSeek-V4-Pro-0813 by DeepSeek ranks #1 on SWE-bench Verified at 96.4%. API pricing is $1.32/M input and $3.96/M output.

    What are the top models on SWE-bench Verified?

    The current SWE-bench Verified ranking as of August 20, 2026 is 1. DeepSeek-V4-Pro-0813 at 96.4%; 2. GPT-5.6 Sol at 96.2%; 3. Claude Opus 5 at 96%.

    Which swe-bench verified model is the cheapest?

    Qwen2.5-Coder 32B Instruct is the cheapest scored model on this swe-bench verified leaderboard at $0.09/M input and $0.09/M output ($0.18 blended). DeepSeek-V4-Pro-0813 still leads SWE-bench Verified at 96.4%.

    Should I always pick the #1 SWE-bench Verified model?

    Not automatically. DeepSeek-V4-Pro-0813 leads SWE-bench Verified, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh SWE-bench Verified against input/output price, context window, and related evals.

    How often is the SWE-bench Verified leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is SWE-bench Verified?

    SWE-bench Verified — a human-validated subset of SWE-bench ensuring each task is solvable and well-specified. This page ranks models that have published a SWE-bench Verified score, with live API token prices on the same row. Official methodology: SWE-bench Verified (https://www.swebench.com/verified.html).

    Where is the SWE-bench Verified leaderboard?

    This page is the SWE-bench Verified leaderboard. Models are sorted by SWE-bench Verified, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this SWE-bench Verified ranking different from the official board?

    The official SWE-bench Verified page owns the methodology. This page keeps the published SWE-bench Verified score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://www.swebench.com/verified.html