Coding

    🏆 Leaderboard

    As of August 22, 2026, DeepSeek-V4-Pro-0813 is #1 for coding at 96.4%. Ranked by the SWE-bench Verified score Best AI for coding, ranked by SWE-bench Verified and Terminal-Bench. Compare Claude vs ChatGPT for coding and see API price on every row.

    Updated August 22, 2026282 models33 providers
    DeepSeek-V4-Pro-0813OSS
    DeepSeek · Open Source
    96.4%87.5%144656.2%82.3%87.9%35.8%41.5%62.7%62.8%$5.28
    GPT-5.6 Sol
    OpenAI · Proprietary
    96.2%82.6%56.9%1620.2780.5%64.6%88.8%86.7%52.9%72.7%72.7%47.5%67.2%$35.00
    Claude Opus 5
    Anthropic · Proprietary
    96%89.0%55.7%1711.8888.4%79.2%84.6%91.7%57.5%73.7%73.7%53.4%70%$30.00
    Grok 4.6
    xAI · Proprietary
    95.6%88.2%163176.2%78.3%44.6%67.5%61.3%$8.00
    Claude Mythos 5
    Anthropic · Proprietary
    95.5%88%80.3%88%80.3%$60.00
    Claude Fable 5
    Anthropic · Proprietary
    95%89.8%60.2%162790.3%80%84.3%72.3%55.1%69.9%69.9%46.3%72.9%80%$60.00
    Claude Mythos Preview
    Anthropic · Proprietary
    93.9%82%77.8%$60.00
    Kimi K3OSS
    Moonshot AI · Open Source
    93.4%87.2%58.7%1681.7588.3%85.0%88.3%16.1%67.5%68.5%44.2%$18.00
    GPT-5.6 Luna
    OpenAI · Proprietary
    93%52.5%1522.9477.1%62.7%84.7%72.9%44.5%67.2%67.2%39.8%$1.40
    DeepSeek-V4-Flash-0731OSS
    DeepSeek · Open Source
    88.8%87.3%49.9%74.7%82.7%38.6%54.4%53.3%$1.76
    Claude Opus 4.8
    Anthropic · Proprietary
    88.6%87.8%53.5%153970.0%82.7%69.2%71.9%47.3%58.2%59.0%46.5%63.8%$30.00
    Claude Opus 4.7
    Anthropic · Proprietary
    87.6%85.1%54.5%155868.5%80.2%71%64.3%68.5%47.1%43.9%54.2%38.5%64.8%$30.00
    Muse Spark 1.2
    Meta · Proprietary
    86.6%153579.1%82.9%49.5%29.9%54.9%$5.50
    Grok 4.5
    xAI · Proprietary
    86.6%87.3%54.0%155583.3%69%64.7%83.3%36.6%53%53.8%42.4%$8.00
    Qwen3.8 MaxOSS
    Qwen · Open Source
    85.6%87.8%166786.6%64.7%67.7%86.6%73%24.0%57.5%$8.00
    Claude Sonnet 5
    Anthropic · Proprietary
    85.2%82.4%53.6%1544.1680.4%80.4%81.3%63.2%74.5%44.4%53.9%53.9%38.8%61.2%$12.00
    GPT-5.5
    OpenAI · Proprietary
    82.6%85.3%56.1%1504.7482.7%84.7%69.8%58.6%76.4%45.2%70.0%67.0%43%64.3%58.6%$35.00
    Muse Spark 1.1
    Meta · Proprietary
    82%85.9%58.2%153872.2%61.5%80%31.1%53.3%53.3%$5.50
    Claude Opus 4.5
    Anthropic · Proprietary
    80.9%75.0%146858.4%63.1%23.6%$30.00
    Gemini 3.7 Flash
    Google · Proprietary
    80.8%88.7%158770.4%85.8%34.8%65.5%$4.50
    Claude Opus 4.6
    Anthropic · Proprietary
    80.8%153765.4%79.8%57.6%27.6%$30.00
    Gemini 3.1 Pro
    Google · Proprietary
    80.6%88.5%59%144767.4%80.2%32.0%54.2%70.8%17.3%9.7%11.7%$17.50
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    80.6%93.5%50%67.9%55.4%$5.22
    MiniMax M3OSS
    MiniMax · Open Source
    80.5%82.2%45.4%149046.1%47.6%59%66%19.9%20.4%14.7%$1.50
    Qwen3.7 Max
    Qwen · Proprietary
    80.4%87.1%53.5%151759.2%47.7%60.6%61.0%46.8%91.6%13.1%17.7%$5.00
    MiniMax M2.5OSS
    MiniMax · Open Source
    80.2%79.2%138441.6%42.7%14.8%55.4%6.7%$1.50
    Kimi K2.6OSS
    Moonshot AI · Open Source
    80.2%86.8%52.2%150957.3%37.9%58.6%53.6%89.6%27.8%23.9%47.6%$4.93
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    80.2%85.9%48.7%19.1%55.9%64.7%13.7%$1.50
    GPT-5.2
    OpenAI · Proprietary
    80%85.4%141851.7%64.9%53.5%79.9%54.8%$15.75
    Gemini 3.6 Flash
    Google · Proprietary
    79.6%88.1%52.7%1527.8378%64.0%58.7%78%30.9%48.6%46.7%58.7%$4.50
    Claude Sonnet 4.6
    Anthropic · Proprietary
    79.6%82.1%46.8%152459.5%53.4%51.5%57.3%39.9%31.8%29.9%49%$18.00
    Gemini 3.5 Flash
    Google · Proprietary
    79.3%87.6%53.1%1506.3867.4%76.2%48.7%55.1%76.2%26.8%28.3%36.1%49.8%$10.50
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    79%91.6%44.9%56.9%52.6%$0.42
    MiMo-V2.5-ProOSS
    Xiaomi · Open Source
    78.9%81.3%50.2%147468.4%34.1%57.2%57.3%75.6%39.6%21.6%19.5%$1.30
    Qwen3.6 Plus
    Qwen · Proprietary
    78.8%86.0%40.7%146044.9%25.6%56.6%53.2%87.1%11.1%2.6%$3.50
    GLM-5.2OSS
    Z AI · Open Source
    78.7%69.5%50.5%1593.2564.0%62.1%82.7%37.9%41.5%43.8%24.5%54.6%$5.80
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    78.6%88.4%56.6%52.3%$0.30
    Kimi K2.7 CodeOSS
    Moonshot AI · Open Source
    78.2%82.0%47.5%147347.2%67.0%25.4%30.5%30.5%30.1%$4.93
    Hy3OSS
    Tencent · Open Source
    78%152257.9%71.7%28%$0.66
    Gemini 3 Flash
    Google · Proprietary
    78%85.6%143851.7%64.3%20.2%53.9%39.1%6.4%5.1%$3.50
    GLM-5OSS
    Z AI · Open Source
    77.8%143656.2%52.4%$4.20
    Qwen3.7-Plus
    Qwen · Proprietary
    77.7%51.3%70.3%46.4%57.6%52.8%89.6%12.9%10.2%$1.60
    Mistral Medium 3.5OSS
    Mistral · Open Source · via Mistral AI
    77.6%39.6%126530.3%2.9%39.0%5.1%$9.00
    Qwen3.6-27BOSS
    Qwen · Open Source
    77.2%44.9%11.9%53.5%83.9%$4.20
    GPT-5.4
    OpenAI · Proprietary
    76.9%84.1%56.6%139062.2%81.8%67.4%57.7%67.8%35.0%55.5%51.8%88.4%$17.50
    Kimi K2.5OSS
    Moonshot AI · Open Source
    76.8%48.7%1430.5650.8%43.2%50.7%85%31.9%$3.68
    Seed 2.0 Pro
    ByteDance · Proprietary
    76.5%87.8%$3.50
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    76.4%140052.5%83.6%$4.20
    GPT-5.1 Thinking
    OpenAI · Proprietary
    76.3%$11.25
    GPT-5.1 Instant
    OpenAI · Proprietary
    76.3%$11.25
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    What is the best AI for coding?

    As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek is #1 for coding at 96.4%. Ranked by the SWE-bench Verified score This board also tracks SWE-bench Verified, LiveCodeBench, SciCode, Arena Code Elo. Next on the same board: GPT-5.6 Sol and Claude Opus 5. Related leaders: DeepSeek-V4-Pro-Max on LiveCodeBench at 93.5%; Claude Fable 5 on SciCode at 60.2%. This coding leaderboard ranks models by SWE-bench Verified. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: SWE-bench Verified (https://www.swebench.com/verified.html); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for coding. Ranked by the SWE-bench Verified score Input and output are dollars per million tokens.
    RankModelSWE-bench VerifiedInput /MOutput /M
    1DeepSeek-V4-Pro-081396.4%$1.32$3.96
    2GPT-5.6 Sol96.2%$5.00$30.00
    3Claude Opus 596%$5.00$25.00
    4Grok 4.695.6%$2.00$6.00
    5Claude Mythos 595.5%$10.00$50.00
    6Claude Fable 595%$10.00$50.00
    7Claude Mythos Preview93.9%$10.00$50.00
    8Kimi K393.4%$3.00$15.00

    What is SWE-bench Verified?

    SWE-bench Verified is 500 real GitHub issues, human-filtered so the tests are fair. Labs quote it because HumanEval saturated. A high Verified score means the model can patch a repo, not that it writes pretty snippets. Check Terminal-Bench if your agent lives in the shell.

    SWE-bench vs Terminal-Bench

    SWE-bench is repository issue fixing. Terminal-Bench is terminal and shell work. Claude often looks stronger on repo patches. GPT-class models often look stronger in the terminal. If your product is an agent, sort both columns before you lock a default model.

    Claude vs ChatGPT for coding

    Do not pick a brand. Rank the specific Claude and GPT rows on SWE-bench Verified and Terminal-Bench, then look at $/M. Cursor vs Claude is the wrong comparison: Cursor is the editor, Claude is one of the models it can call.

    Coding FAQ

    Who ranks #1 on the Coding leaderboard?

    As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek ranks #1 on SWE-bench Verified at 96.4%. API pricing is $1.32/M input and $3.96/M output.

    What is the best AI for coding?

    As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek is #1 for coding at 96.4%. Ranked by the SWE-bench Verified score This board also tracks SWE-bench Verified, LiveCodeBench, SciCode, Arena Code Elo. Next on the same board: GPT-5.6 Sol and Claude Opus 5. Related leaders: DeepSeek-V4-Pro-Max on LiveCodeBench at 93.5%; Claude Fable 5 on SciCode at 60.2%.

    What are the top models on SWE-bench Verified?

    The current SWE-bench Verified ranking as of August 22, 2026 is 1. DeepSeek-V4-Pro-0813 at 96.4%; 2. GPT-5.6 Sol at 96.2%; 3. Claude Opus 5 at 96%.

    Which coding model is the cheapest?

    Qwen2.5-Coder 32B Instruct is the cheapest scored model on this coding leaderboard at $0.09/M input and $0.09/M output ($0.18 blended). DeepSeek-V4-Pro-0813 still leads SWE-bench Verified at 96.4%.

    Should I always pick the #1 SWE-bench Verified model?

    Not automatically. DeepSeek-V4-Pro-0813 leads SWE-bench Verified, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh SWE-bench Verified against input/output price, context window, and related evals.

    How often is the Coding leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 22, 2026. Treat it as a current index, not a one-off blog post.

    What is the best LLM for coding right now?

    As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek leads this coding leaderboard on SWE-bench Verified at 96.4%. API pricing is $1.32/M input and $3.96/M output. If a cheaper scored model is close on SWE-bench, that is often the better production pick.

    What is the best AI for coding right now?

    The best coding model depends on the eval. SWE-bench Verified measures real GitHub fixes, LiveCodeBench tracks contest programming, and Terminal-Bench tests shell agents. Rank the table by the job you actually run.

    Is SWE-bench the same as LiveCodeBench?

    No. SWE-bench scores repository-level issue fixing. LiveCodeBench scores fresh competitive-programming problems. A model can lead one and lag the other.

    Do coding leaderboard scores include API price?

    Yes. Every row shows input and output token prices so you can see whether the top SWE-bench model is worth the extra cost versus an open-source coder.

    What is the best coding AI?

    Same question as best AI for coding. Rank SWE-bench Verified for GitHub issue fixes and Terminal-Bench for shell agents. A model can lead one and lag the other. Price sits on the same row so you can skip a tiny score gap that triples the bill.

    What is SWE-bench?

    SWE-bench scores whether a model can resolve real GitHub issues. SWE-bench Verified is the human-filtered 500-issue subset most labs quote. SWE-bench Pro and DeepSWE are harder follow-ons. HumanEval is a short function-completion test, not the same job.

    What is SWE-bench Verified?

    SWE-bench Verified is 500 human-checked GitHub issues. It is the default public coding rank in 2026. Bash-only and full-agent harnesses are not interchangeable, so compare scores that name the same setup.

    What is Terminal-Bench?

    Terminal-Bench scores shell and terminal agents: install tools, run commands, finish a task in a real environment. That is closer to a coding agent loop than HumanEval. SWE-bench is repo patches. Terminal-Bench is the terminal. Sort both.

    Claude vs ChatGPT for coding: which is better?

    Claude vs ChatGPT for coding is a SKU fight. Sonnet/Opus-class Claude and GPT-class OpenAI models trade SWE-bench and Terminal-Bench leads as evals refresh. Rank this table, then take the cheaper model if the gap is small.

    Cursor vs Claude: which is better for coding?

    Cursor is an IDE that calls models, often Claude. Claude is the model. Ranking Cursor vs Claude mixes a product and a model. Pick the model on this SWE-bench board, then pick the client (Cursor, Codex, Claude Code) separately.

    What is the best AI coding assistant?

    An AI coding assistant is the client plus the model. This page ranks the models on SWE-bench Verified, Terminal-Bench, LiveCodeBench, and related evals, with API price. The assistant UI is a different purchase.

    What is the best open source LLM for coding?

    Open the open-source board and sort by SWE-bench or LiveCodeBench. Qwen, DeepSeek, Kimi, and Llama variants rotate. Weights can be free. The hosted coding API still bills tokens.