SWE-bench

    🏆 Leaderboard

    As of August 20, 2026, DeepSeek-V4-Pro-0813 is #1 for SWE-bench at 96.4%. Ranked by the SWE-bench Verified score SWE-bench leaderboard with Verified, Pro, and DeepSWE scores plus API pricing. See which LLM actually fixes real GitHub issues.

    Updated August 20, 2026282 models33 providers
    DeepSeek-V4-Pro-0813OSS
    DeepSeek · Open Source
    96.4%62.7%62.8%$5.28
    GPT-5.6 Sol
    OpenAI · Proprietary
    96.2%64.6%72.7%72.7%$35.00
    Claude Opus 5
    Anthropic · Proprietary
    96%79.2%73.7%73.7%$30.00
    Grok 4.6
    xAI · Proprietary
    95.6%67.5%$8.00
    Claude Mythos 5
    Anthropic · Proprietary
    95.5%80.3%77.8%80.3%$60.00
    Claude Fable 5
    Anthropic · Proprietary
    95%80%69.9%69.9%80%$60.00
    Claude Mythos Preview
    Anthropic · Proprietary
    93.9%77.8%87.3%$60.00
    Kimi K3OSS
    Moonshot AI · Open Source
    93.4%67.5%68.5%$18.00
    GPT-5.6 Luna
    OpenAI · Proprietary
    93%62.7%67.2%67.2%$1.40
    DeepSeek-V4-Flash-0731OSS
    DeepSeek · Open Source
    88.8%54.4%53.3%$1.76
    Claude Opus 4.8
    Anthropic · Proprietary
    88.6%69.2%84.4%58.2%59.0%$30.00
    Claude Opus 4.7
    Anthropic · Proprietary
    87.6%64.3%54.2%$30.00
    Muse Spark 1.2
    Meta · Proprietary
    86.6%54.9%$5.50
    Grok 4.5
    xAI · Proprietary
    86.6%64.7%53%53.8%$8.00
    Qwen3.8 MaxOSS
    Qwen · Open Source
    85.6%67.7%57.5%$8.00
    Claude Sonnet 5
    Anthropic · Proprietary
    85.2%63.2%78.3%53.9%53.9%$12.00
    GPT-5.5
    OpenAI · Proprietary
    82.6%58.6%70.0%67.0%58.6%$35.00
    Muse Spark 1.1
    Meta · Proprietary
    82%61.5%53.3%53.3%$5.50
    Claude Opus 4.5
    Anthropic · Proprietary
    80.9%$30.00
    Gemini 3.7 Flash
    Google · Proprietary
    80.8%65.5%$4.50
    Claude Opus 4.6
    Anthropic · Proprietary
    80.8%77.8%27.6%$30.00
    Gemini 3.1 Pro
    Google · Proprietary
    80.6%54.2%9.7%11.7%$17.50
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    80.6%55.4%76.2%$5.22
    MiniMax M3OSS
    MiniMax · Open Source
    80.5%59%20.4%$1.50
    Qwen3.7 Max
    Qwen · Proprietary
    80.4%60.6%78.3%17.7%$5.00
    MiniMax M2.5OSS
    MiniMax · Open Source
    80.2%55.4%$1.50
    Kimi K2.6OSS
    Moonshot AI · Open Source
    80.2%58.6%76.7%23.9%$4.93
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    80.2%55.9%$1.50
    GPT-5.2
    OpenAI · Proprietary
    80%79.9%$15.75
    Gemini 3.6 Flash
    Google · Proprietary
    79.6%58.7%48.6%46.7%58.7%$4.50
    Claude Sonnet 4.6
    Anthropic · Proprietary
    79.6%31.8%29.9%$18.00
    Gemini 3.5 Flash
    Google · Proprietary
    79.3%55.1%28.3%36.1%$10.50
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    79%52.6%73.3%$0.42
    MiMo-V2.5-ProOSS
    Xiaomi · Open Source
    78.9%57.2%19.5%$1.30
    Qwen3.6 Plus
    Qwen · Proprietary
    78.8%56.6%73.8%2.6%$3.50
    GLM-5.2OSS
    Z AI · Open Source
    78.7%62.1%41.5%43.8%$5.80
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    78.6%52.3%70.2%$0.30
    Kimi K2.7 CodeOSS
    Moonshot AI · Open Source
    78.2%30.5%30.5%$4.93
    Hy3OSS
    Tencent · Open Source
    78%57.9%75.8%28%$0.66
    Gemini 3 Flash
    Google · Proprietary
    78%5.1%$3.50
    GLM-5OSS
    Z AI · Open Source
    77.8%$4.20
    Qwen3.7-Plus
    Qwen · Proprietary
    77.7%57.6%75.8%$1.60
    Mistral Medium 3.5OSS
    Mistral · Open Source · via Mistral AI
    77.6%$9.00
    Qwen3.6-27BOSS
    Qwen · Open Source
    77.2%53.5%71.3%$4.20
    GPT-5.4
    OpenAI · Proprietary
    76.9%57.7%55.5%51.8%88.4%$17.50
    Kimi K2.5OSS
    Moonshot AI · Open Source
    76.8%50.7%73%$3.68
    Seed 2.0 Pro
    ByteDance · Proprietary
    76.5%$3.50
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    76.4%69.3%$4.20
    GPT-5.1 Thinking
    OpenAI · Proprietary
    76.3%$11.25
    GPT-5.1 Instant
    OpenAI · Proprietary
    76.3%$11.25
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads SWE-bench right now?

    As of August 20, 2026, DeepSeek-V4-Pro-0813 by DeepSeek is #1 for SWE-bench at 96.4%. Ranked by the SWE-bench Verified score This board also tracks SWE-bench Verified, SWE-bench Pro, SWE-bench Multilingual, DeepSWE. Next on the same board: GPT-5.6 Sol and Claude Opus 5. Related leaders: Gemini 3 Pro on SWE-bench Pro at 84.8%; Claude Mythos Preview on SWE-bench Multilingual at 87.3%. This swe-bench leaderboard ranks models by SWE-bench Verified. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: SWE-bench Verified (https://www.swebench.com/verified.html); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for swe-bench. Ranked by the SWE-bench Verified score Input and output are dollars per million tokens.
    RankModelSWE-bench VerifiedInput /MOutput /M
    1DeepSeek-V4-Pro-081396.4%$1.32$3.96
    2GPT-5.6 Sol96.2%$5.00$30.00
    3Claude Opus 596%$5.00$25.00
    4Grok 4.695.6%$2.00$6.00
    5Claude Mythos 595.5%$10.00$50.00
    6Claude Fable 595%$10.00$50.00
    7Claude Mythos Preview93.9%$10.00$50.00
    8Kimi K393.4%$3.00$15.00

    SWE-bench FAQ

    Who ranks #1 on the SWE-bench leaderboard?

    As of August 20, 2026, DeepSeek-V4-Pro-0813 by DeepSeek ranks #1 on SWE-bench Verified at 96.4%. API pricing is $1.32/M input and $3.96/M output.

    What are the top models on SWE-bench Verified?

    The current SWE-bench Verified ranking as of August 20, 2026 is 1. DeepSeek-V4-Pro-0813 at 96.4%; 2. GPT-5.6 Sol at 96.2%; 3. Claude Opus 5 at 96%.

    Which swe-bench model is the cheapest?

    Qwen2.5-Coder 32B Instruct is the cheapest scored model on this swe-bench leaderboard at $0.09/M input and $0.09/M output ($0.18 blended). DeepSeek-V4-Pro-0813 still leads SWE-bench Verified at 96.4%.

    Should I always pick the #1 SWE-bench Verified model?

    Not automatically. DeepSeek-V4-Pro-0813 leads SWE-bench Verified, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh SWE-bench Verified against input/output price, context window, and related evals.

    How often is the SWE-bench leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is SWE-bench?

    SWE-bench asks a model to resolve real GitHub issues by editing a repository until the tests pass. SWE-bench Verified is the human-checked subset.

    What is the difference between SWE-bench Verified and SWE-bench Pro?

    Verified is the standard public subset. Pro is a harder set of longer, messier engineering tasks. DeepSWE is a separate agent-style software eval.

    Where can I see SWE-bench with prices?

    Here. Official SWE-bench sites rank accuracy only. This table adds input and output token prices so you can compare a 70% model at $1/M versus a 75% model at $15/M.