SWE-bench Pro

    🏆 Leaderboard

    As of August 20, 2026, Gemini 3 Pro is #1 for SWE-bench Pro at 84.8%. Ranked by the SWE-bench Pro score 52 models in this index have a published SWE-bench Pro score. Methodology: SWE-bench (https://www.swebench.com/). SWE-bench Pro leaderboard with live API prices. SWE-bench Pro — harder real-world GitHub issue solving than Verified. Official methodology: SWE-bench (https://www.swebench.com/).

    Updated August 20, 2026282 models33 providers
    Gemini 3 Pro
    Google · Proprietary
    84.8%$14.00
    GPT-5.2 Pro
    OpenAI · Proprietary
    84.5%$189.00
    Claude Mythos 5
    Anthropic · Proprietary
    80.3%$60.00
    Claude Fable 5
    Anthropic · Proprietary
    80%$60.00
    GPT-5.2
    OpenAI · Proprietary
    79.9%$15.75
    Claude Opus 5
    Anthropic · Proprietary
    79.2%$30.00
    Claude Mythos Preview
    Anthropic · Proprietary
    77.8%$60.00
    Claude Sonnet 4
    Anthropic · Proprietary
    74.8%$18.00
    Claude Opus 4.8
    Anthropic · Proprietary
    69.2%$30.00
    Qwen3.8 MaxOSS
    Qwen · Open Source
    67.7%$8.00
    Grok 4.5
    xAI · Proprietary
    64.7%$8.00
    GPT-5.6 Sol
    OpenAI · Proprietary
    64.6%$35.00
    Claude Opus 4.7
    Anthropic · Proprietary
    64.3%$30.00
    GPT-5.6 Terra
    OpenAI · Proprietary
    63.4%$14.00
    Claude Sonnet 5
    Anthropic · Proprietary
    63.2%$12.00
    GPT-5.6 Luna
    OpenAI · Proprietary
    62.7%$1.40
    GLM-5.2OSS
    Z AI · Open Source
    62.1%$5.80
    Qwen3.8-27BOSS
    Qwen · Open Source
    61.7%$3.65
    Muse Spark 1.1
    Meta · Proprietary
    61.5%$5.50
    Qwen3.7 Max
    Qwen · Proprietary
    60.6%$5.00
    Laguna S 2.1OSS
    Poolside · Open Source
    59.4%$0.30
    MiniMax M3OSS
    MiniMax · Open Source
    59%$1.50
    Gemini 3.6 Flash
    Google · Proprietary
    58.7%$4.50
    Kimi K2.6OSS
    Moonshot AI · Open Source
    58.6%$4.93
    GPT-5.5
    OpenAI · Proprietary
    58.6%$35.00
    GLM-5.1OSS
    Z AI · Open Source
    58.4%$5.80
    Hy3OSS
    Tencent · Open Source
    57.9%$0.66
    GPT-5.4
    OpenAI · Proprietary
    57.7%$17.50
    Qwen3.7-Plus
    Qwen · Proprietary
    57.6%$1.60
    MiMo-V2.5-ProOSS
    Xiaomi · Open Source
    57.2%$1.30
    Seed 2.1 Turbo
    ByteDance · Proprietary
    57%$3.00
    GPT-5.3 Codex
    OpenAI · Proprietary
    56.8%$15.75
    Qwen3.6 Plus
    Qwen · Proprietary
    56.6%$3.50
    GPT-5.2 Codex
    OpenAI · Proprietary
    56.4%$15.75
    MiniMax M2.7OSS
    MiniMax · Open Source
    56.2%$1.50
    MiMo-V2.5OSS
    Xiaomi · Open Source
    56.1%$0.50
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    55.9%$1.50
    MiniMax M2.5OSS
    MiniMax · Open Source
    55.4%$1.50
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    55.4%$5.22
    Gemini 3.5 Flash
    Google · Proprietary
    55.1%$10.50
    GPT-5.4 mini
    OpenAI · Proprietary
    54.4%$5.25
    Gemini 3.5 Flash-Lite
    Google · Proprietary
    54.2%$2.80
    Gemini 3.1 Pro
    Google · Proprietary
    54.2%$17.50
    Qwen3.6-27BOSS
    Qwen · Open Source
    53.5%$4.20
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    52.6%$0.42
    GPT-5.4 nano
    OpenAI · Proprietary
    52.4%$1.45
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    52.3%$0.30
    Muse Glimmer-30BOSS
    Meta · Open Source
    51.2%$1.85
    Kimi K2.5OSS
    Moonshot AI · Open Source
    50.7%$3.68
    Qwen3.6-35B-A3BOSS
    Qwen · Open Source · via OpenRouter
    49.5%$1.14
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads SWE-bench Pro right now?

    As of August 20, 2026, Gemini 3 Pro by Google is #1 for SWE-bench Pro at 84.8%. Ranked by the SWE-bench Pro score This board also tracks SWE-bench Pro. Next on the same board: GPT-5.2 Pro and Claude Mythos 5. This swe-bench pro leaderboard ranks models by SWE-bench Pro. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: SWE-bench (https://www.swebench.com/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for swe-bench pro. Ranked by the SWE-bench Pro score Input and output are dollars per million tokens.
    RankModelSWE-bench ProInput /MOutput /M
    1Gemini 3 Pro84.8%$2.00$12.00
    2GPT-5.2 Pro84.5%$21.00$168.00
    3Claude Mythos 580.3%$10.00$50.00
    4Claude Fable 580%$10.00$50.00
    5GPT-5.279.9%$1.75$14.00
    6Claude Opus 579.2%$5.00$25.00
    7Claude Mythos Preview77.8%$10.00$50.00
    8Claude Sonnet 474.8%$3.00$15.00

    SWE-bench Pro FAQ

    Who ranks #1 on the SWE-bench Pro leaderboard?

    As of August 20, 2026, Gemini 3 Pro by Google ranks #1 on SWE-bench Pro at 84.8%. API pricing is $2.00/M input and $12.00/M output.

    What are the top models on SWE-bench Pro?

    The current SWE-bench Pro ranking as of August 20, 2026 is 1. Gemini 3 Pro at 84.8%; 2. GPT-5.2 Pro at 84.5%; 3. Claude Mythos 5 at 80.3%.

    Which swe-bench pro model is the cheapest?

    Laguna S 2.1 is the cheapest scored model on this swe-bench pro leaderboard at $0.10/M input and $0.20/M output ($0.30 blended). Gemini 3 Pro still leads SWE-bench Pro at 84.8%.

    Should I always pick the #1 SWE-bench Pro model?

    Not automatically. Gemini 3 Pro leads SWE-bench Pro, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh SWE-bench Pro against input/output price, context window, and related evals.

    How often is the SWE-bench Pro leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is SWE-bench Pro?

    SWE-bench Pro — harder real-world GitHub issue solving than Verified. This page ranks models that have published a SWE-bench Pro score, with live API token prices on the same row. Official methodology: SWE-bench (https://www.swebench.com/).

    Where is the SWE-bench Pro leaderboard?

    This page is the SWE-bench Pro leaderboard. Models are sorted by SWE-bench Pro, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this SWE-bench Pro ranking different from the official board?

    The official SWE-bench page owns the methodology. This page keeps the published SWE-bench Pro score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://www.swebench.com/