Who ranks #1 on the Coding leaderboard?
As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek ranks #1 on SWE-bench Verified at 96.4%. API pricing is $1.32/M input and $3.96/M output.
As of August 22, 2026, DeepSeek-V4-Pro-0813 is #1 for coding at 96.4%. Ranked by the SWE-bench Verified score Best AI for coding, ranked by SWE-bench Verified and Terminal-Bench. Compare Claude vs ChatGPT for coding and see API price on every row.
DeepSeek-V4-Pro-0813OSS DeepSeek · Open Source | 96.4% | 87.5% | — | 1446 | 56.2% | — | 82.3% | — | 87.9% | — | 35.8% | — | — | 41.5% | 62.7% | 62.8% | — | — | — | $5.28 |
![]() GPT-5.6 Sol OpenAI · Proprietary | 96.2% | 82.6% | 56.9% | 1620.27 | — | — | 80.5% | 64.6% | 88.8% | — | 86.7% | — | — | 52.9% | 72.7% | 72.7% | 47.5% | 67.2% | — | $35.00 |
![]() Claude Opus 5 Anthropic · Proprietary | 96% | 89.0% | 55.7% | 1711.88 | — | — | 88.4% | 79.2% | 84.6% | — | 91.7% | — | — | 57.5% | 73.7% | 73.7% | 53.4% | 70% | — | $30.00 |
![]() Grok 4.6 xAI · Proprietary | 95.6% | 88.2% | — | 1631 | — | — | 76.2% | — | 78.3% | — | — | — | — | 44.6% | — | 67.5% | 61.3% | — | — | $8.00 |
![]() Claude Mythos 5 Anthropic · Proprietary | 95.5% | — | — | — | — | 88% | — | 80.3% | 88% | — | — | — | — | — | — | — | — | — | 80.3% | $60.00 |
![]() Claude Fable 5 Anthropic · Proprietary | 95% | 89.8% | 60.2% | 1627 | — | — | 90.3% | 80% | 84.3% | — | 72.3% | — | — | 55.1% | 69.9% | 69.9% | 46.3% | 72.9% | 80% | $60.00 |
![]() Claude Mythos Preview Anthropic · Proprietary | 93.9% | — | — | — | 82% | — | — | 77.8% | — | — | — | — | — | — | — | — | — | — | — | $60.00 |
Kimi K3OSS Moonshot AI · Open Source | 93.4% | 87.2% | 58.7% | 1681.75 | — | 88.3% | 85.0% | — | 88.3% | — | — | — | — | 16.1% | 67.5% | 68.5% | 44.2% | — | — | $18.00 |
![]() GPT-5.6 Luna OpenAI · Proprietary | 93% | — | 52.5% | 1522.94 | — | — | 77.1% | 62.7% | 84.7% | — | 72.9% | — | — | 44.5% | 67.2% | 67.2% | 39.8% | — | — | $1.40 |
DeepSeek-V4-Flash-0731OSS DeepSeek · Open Source | 88.8% | 87.3% | 49.9% | — | — | — | 74.7% | — | 82.7% | — | — | — | — | 38.6% | 54.4% | 53.3% | — | — | — | $1.76 |
![]() Claude Opus 4.8 Anthropic · Proprietary | 88.6% | 87.8% | 53.5% | 1539 | 70.0% | — | 82.7% | 69.2% | 71.9% | — | — | — | — | 47.3% | 58.2% | 59.0% | 46.5% | 63.8% | — | $30.00 |
![]() Claude Opus 4.7 Anthropic · Proprietary | 87.6% | 85.1% | 54.5% | 1558 | 68.5% | 80.2% | 71% | 64.3% | 68.5% | — | 47.1% | — | — | 43.9% | 54.2% | — | 38.5% | 64.8% | — | $30.00 |
Muse Spark 1.2 Meta · Proprietary | 86.6% | — | — | 1535 | — | — | 79.1% | — | 82.9% | — | 49.5% | — | — | 29.9% | — | 54.9% | — | — | — | $5.50 |
![]() Grok 4.5 xAI · Proprietary | 86.6% | 87.3% | 54.0% | 1555 | — | 83.3% | 69% | 64.7% | 83.3% | — | — | — | — | 36.6% | 53% | 53.8% | 42.4% | — | — | $8.00 |
Qwen3.8 MaxOSS Qwen · Open Source | 85.6% | 87.8% | — | 1667 | — | 86.6% | 64.7% | 67.7% | 86.6% | — | 73% | — | — | 24.0% | — | 57.5% | — | — | — | $8.00 |
![]() Claude Sonnet 5 Anthropic · Proprietary | 85.2% | 82.4% | 53.6% | 1544.16 | 80.4% | 80.4% | 81.3% | 63.2% | 74.5% | — | — | — | — | 44.4% | 53.9% | 53.9% | 38.8% | 61.2% | — | $12.00 |
![]() GPT-5.5 OpenAI · Proprietary | 82.6% | 85.3% | 56.1% | 1504.74 | 82.7% | 84.7% | 69.8% | 58.6% | 76.4% | — | — | — | — | 45.2% | 70.0% | 67.0% | 43% | 64.3% | 58.6% | $35.00 |
Muse Spark 1.1 Meta · Proprietary | 82% | 85.9% | 58.2% | 1538 | — | — | 72.2% | 61.5% | 80% | — | — | — | — | 31.1% | 53.3% | 53.3% | — | — | — | $5.50 |
![]() Claude Opus 4.5 Anthropic · Proprietary | 80.9% | 75.0% | — | 1468 | 58.4% | 63.1% | — | — | — | — | 23.6% | — | — | — | — | — | — | — | — | $30.00 |
Gemini 3.7 Flash Google · Proprietary | 80.8% | 88.7% | — | 1587 | — | — | 70.4% | — | 85.8% | — | — | — | — | 34.8% | — | 65.5% | — | — | — | $4.50 |
![]() Claude Opus 4.6 Anthropic · Proprietary | 80.8% | — | — | 1537 | 65.4% | 79.8% | 57.6% | — | — | — | — | — | — | — | 27.6% | — | — | — | — | $30.00 |
Gemini 3.1 Pro Google · Proprietary | 80.6% | 88.5% | 59% | 1447 | 67.4% | 80.2% | 32.0% | 54.2% | 70.8% | — | — | — | — | 17.3% | 9.7% | 11.7% | — | — | — | $17.50 |
DeepSeek-V4-Pro-MaxOSS DeepSeek · Open Source | 80.6% | 93.5% | 50% | — | 67.9% | — | — | 55.4% | — | — | — | — | — | — | — | — | — | — | — | $5.22 |
MiniMax M3OSS MiniMax · Open Source | 80.5% | 82.2% | 45.4% | 1490 | 46.1% | — | 47.6% | 59% | 66% | — | — | — | — | 19.9% | 20.4% | — | 14.7% | — | — | $1.50 |
Qwen3.7 Max Qwen · Proprietary | 80.4% | 87.1% | 53.5% | 1517 | 59.2% | — | 47.7% | 60.6% | 61.0% | — | 46.8% | 91.6% | — | 13.1% | 17.7% | — | — | — | — | $5.00 |
MiniMax M2.5OSS MiniMax · Open Source | 80.2% | 79.2% | — | 1384 | 41.6% | 42.7% | 14.8% | 55.4% | — | — | 6.7% | — | — | — | — | — | — | — | — | $1.50 |
Kimi K2.6OSS Moonshot AI · Open Source | 80.2% | 86.8% | 52.2% | 1509 | 57.3% | — | 37.9% | 58.6% | 53.6% | — | — | 89.6% | — | 27.8% | 23.9% | — | — | 47.6% | — | $4.93 |
Inkling-SmallOSS Thinking Machines · Open Source · via Thinking Machines Lab | 80.2% | 85.9% | 48.7% | — | — | — | 19.1% | 55.9% | 64.7% | — | — | — | — | 13.7% | — | — | — | — | — | $1.50 |
![]() GPT-5.2 OpenAI · Proprietary | 80% | 85.4% | — | 1418 | 51.7% | 64.9% | 53.5% | 79.9% | — | — | 54.8% | — | — | — | — | — | — | — | — | $15.75 |
Gemini 3.6 Flash Google · Proprietary | 79.6% | 88.1% | 52.7% | 1527.83 | — | 78% | 64.0% | 58.7% | 78% | — | — | — | — | 30.9% | 48.6% | 46.7% | — | — | 58.7% | $4.50 |
![]() Claude Sonnet 4.6 Anthropic · Proprietary | 79.6% | 82.1% | 46.8% | 1524 | 59.5% | 53.4% | 51.5% | — | 57.3% | — | — | — | — | 39.9% | 31.8% | 29.9% | — | 49% | — | $18.00 |
Gemini 3.5 Flash Google · Proprietary | 79.3% | 87.6% | 53.1% | 1506.38 | 67.4% | 76.2% | 48.7% | 55.1% | 76.2% | — | — | — | — | 26.8% | 28.3% | 36.1% | — | 49.8% | — | $10.50 |
DeepSeek-V4-Flash-MaxOSS DeepSeek · Open Source | 79% | 91.6% | 44.9% | — | 56.9% | — | — | 52.6% | — | — | — | — | — | — | — | — | — | — | — | $0.42 |
MiMo-V2.5-ProOSS Xiaomi · Open Source | 78.9% | 81.3% | 50.2% | 1474 | 68.4% | — | 34.1% | 57.2% | 57.3% | 75.6% | — | 39.6% | — | 21.6% | 19.5% | — | — | — | — | $1.30 |
Qwen3.6 Plus Qwen · Proprietary | 78.8% | 86.0% | 40.7% | 1460 | 44.9% | — | 25.6% | 56.6% | 53.2% | — | — | 87.1% | — | 11.1% | 2.6% | — | — | — | — | $3.50 |
GLM-5.2OSS Z AI · Open Source | 78.7% | 69.5% | 50.5% | 1593.25 | — | — | 64.0% | 62.1% | 82.7% | — | — | — | — | 37.9% | 41.5% | 43.8% | 24.5% | 54.6% | — | $5.80 |
DeepSeek-V4-Flash-0423OSS DeepSeek · Open Source | 78.6% | 88.4% | — | — | 56.6% | — | — | 52.3% | — | — | — | — | — | — | — | — | — | — | — | $0.30 |
Kimi K2.7 CodeOSS Moonshot AI · Open Source | 78.2% | 82.0% | 47.5% | 1473 | — | — | 47.2% | — | 67.0% | — | — | — | — | 25.4% | 30.5% | 30.5% | 30.1% | — | — | $4.93 |
Hy3OSS Tencent · Open Source | 78% | — | — | 1522 | — | — | — | 57.9% | 71.7% | — | — | — | — | — | 28% | — | — | — | — | $0.66 |
Gemini 3 Flash Google · Proprietary | 78% | 85.6% | — | 1438 | 51.7% | 64.3% | 20.2% | — | 53.9% | — | 39.1% | — | — | 6.4% | 5.1% | — | — | — | — | $3.50 |
GLM-5OSS Z AI · Open Source | 77.8% | — | — | 1436 | 56.2% | 52.4% | — | — | — | — | — | — | — | — | — | — | — | — | — | $4.20 |
Qwen3.7-Plus Qwen · Proprietary | 77.7% | — | 51.3% | — | 70.3% | — | 46.4% | 57.6% | 52.8% | — | — | 89.6% | — | 12.9% | — | — | 10.2% | — | — | $1.60 |
Mistral Medium 3.5OSS Mistral · Open Source · via Mistral AI | 77.6% | — | 39.6% | 1265 | 30.3% | — | 2.9% | — | 39.0% | — | — | — | — | 5.1% | — | — | — | — | — | $9.00 |
Qwen3.6-27BOSS Qwen · Open Source | 77.2% | — | — | — | 44.9% | — | 11.9% | 53.5% | — | — | — | 83.9% | — | — | — | — | — | — | — | $4.20 |
![]() GPT-5.4 OpenAI · Proprietary | 76.9% | 84.1% | 56.6% | 1390 | 62.2% | 81.8% | 67.4% | 57.7% | — | — | 67.8% | — | — | 35.0% | 55.5% | 51.8% | — | — | 88.4% | $17.50 |
Kimi K2.5OSS Moonshot AI · Open Source | 76.8% | — | 48.7% | 1430.56 | 50.8% | 43.2% | — | 50.7% | — | — | — | 85% | — | — | — | — | — | 31.9% | — | $3.68 |
Seed 2.0 Pro ByteDance · Proprietary | 76.5% | — | — | — | — | — | — | — | — | — | — | 87.8% | — | — | — | — | — | — | — | $3.50 |
Qwen3.5-397B-A17BOSS Qwen · Open Source | 76.4% | — | — | 1400 | 52.5% | — | — | — | — | — | — | 83.6% | — | — | — | — | — | — | — | $4.20 |
![]() GPT-5.1 Thinking OpenAI · Proprietary | 76.3% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | $11.25 |
![]() GPT-5.1 Instant OpenAI · Proprietary | 76.3% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | $11.25 |
Next step
Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.
The index
282 models across 33 providers. Search or jump to a lab — every model page stays linked here.







As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek is #1 for coding at 96.4%. Ranked by the SWE-bench Verified score This board also tracks SWE-bench Verified, LiveCodeBench, SciCode, Arena Code Elo. Next on the same board: GPT-5.6 Sol and Claude Opus 5. Related leaders: DeepSeek-V4-Pro-Max on LiveCodeBench at 93.5%; Claude Fable 5 on SciCode at 60.2%. This coding leaderboard ranks models by SWE-bench Verified. Scores come from public evals. Prices are the live API rates in the table above.
Sources: SWE-bench Verified (https://www.swebench.com/verified.html); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)
| Rank | Model | SWE-bench Verified | Input /M | Output /M |
|---|---|---|---|---|
| 1 | DeepSeek-V4-Pro-0813 | 96.4% | $1.32 | $3.96 |
| 2 | GPT-5.6 Sol | 96.2% | $5.00 | $30.00 |
| 3 | Claude Opus 5 | 96% | $5.00 | $25.00 |
| 4 | Grok 4.6 | 95.6% | $2.00 | $6.00 |
| 5 | Claude Mythos 5 | 95.5% | $10.00 | $50.00 |
| 6 | Claude Fable 5 | 95% | $10.00 | $50.00 |
| 7 | Claude Mythos Preview | 93.9% | $10.00 | $50.00 |
| 8 | Kimi K3 | 93.4% | $3.00 | $15.00 |
SWE-bench Verified is 500 real GitHub issues, human-filtered so the tests are fair. Labs quote it because HumanEval saturated. A high Verified score means the model can patch a repo, not that it writes pretty snippets. Check Terminal-Bench if your agent lives in the shell.
SWE-bench is repository issue fixing. Terminal-Bench is terminal and shell work. Claude often looks stronger on repo patches. GPT-class models often look stronger in the terminal. If your product is an agent, sort both columns before you lock a default model.
Do not pick a brand. Rank the specific Claude and GPT rows on SWE-bench Verified and Terminal-Bench, then look at $/M. Cursor vs Claude is the wrong comparison: Cursor is the editor, Claude is one of the models it can call.
Rank one eval at a time. All LLM benchmarks.
As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek ranks #1 on SWE-bench Verified at 96.4%. API pricing is $1.32/M input and $3.96/M output.
As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek is #1 for coding at 96.4%. Ranked by the SWE-bench Verified score This board also tracks SWE-bench Verified, LiveCodeBench, SciCode, Arena Code Elo. Next on the same board: GPT-5.6 Sol and Claude Opus 5. Related leaders: DeepSeek-V4-Pro-Max on LiveCodeBench at 93.5%; Claude Fable 5 on SciCode at 60.2%.
The current SWE-bench Verified ranking as of August 22, 2026 is 1. DeepSeek-V4-Pro-0813 at 96.4%; 2. GPT-5.6 Sol at 96.2%; 3. Claude Opus 5 at 96%.
Qwen2.5-Coder 32B Instruct is the cheapest scored model on this coding leaderboard at $0.09/M input and $0.09/M output ($0.18 blended). DeepSeek-V4-Pro-0813 still leads SWE-bench Verified at 96.4%.
Not automatically. DeepSeek-V4-Pro-0813 leads SWE-bench Verified, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh SWE-bench Verified against input/output price, context window, and related evals.
Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 22, 2026. Treat it as a current index, not a one-off blog post.
As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek leads this coding leaderboard on SWE-bench Verified at 96.4%. API pricing is $1.32/M input and $3.96/M output. If a cheaper scored model is close on SWE-bench, that is often the better production pick.
The best coding model depends on the eval. SWE-bench Verified measures real GitHub fixes, LiveCodeBench tracks contest programming, and Terminal-Bench tests shell agents. Rank the table by the job you actually run.
No. SWE-bench scores repository-level issue fixing. LiveCodeBench scores fresh competitive-programming problems. A model can lead one and lag the other.
Yes. Every row shows input and output token prices so you can see whether the top SWE-bench model is worth the extra cost versus an open-source coder.
Same question as best AI for coding. Rank SWE-bench Verified for GitHub issue fixes and Terminal-Bench for shell agents. A model can lead one and lag the other. Price sits on the same row so you can skip a tiny score gap that triples the bill.
SWE-bench scores whether a model can resolve real GitHub issues. SWE-bench Verified is the human-filtered 500-issue subset most labs quote. SWE-bench Pro and DeepSWE are harder follow-ons. HumanEval is a short function-completion test, not the same job.
SWE-bench Verified is 500 human-checked GitHub issues. It is the default public coding rank in 2026. Bash-only and full-agent harnesses are not interchangeable, so compare scores that name the same setup.
Terminal-Bench scores shell and terminal agents: install tools, run commands, finish a task in a real environment. That is closer to a coding agent loop than HumanEval. SWE-bench is repo patches. Terminal-Bench is the terminal. Sort both.
Claude vs ChatGPT for coding is a SKU fight. Sonnet/Opus-class Claude and GPT-class OpenAI models trade SWE-bench and Terminal-Bench leads as evals refresh. Rank this table, then take the cheaper model if the gap is small.
Cursor is an IDE that calls models, often Claude. Claude is the model. Ranking Cursor vs Claude mixes a product and a model. Pick the model on this SWE-bench board, then pick the client (Cursor, Codex, Claude Code) separately.
An AI coding assistant is the client plus the model. This page ranks the models on SWE-bench Verified, Terminal-Bench, LiveCodeBench, and related evals, with API price. The assistant UI is a different purchase.
Open the open-source board and sort by SWE-bench or LiveCodeBench. Qwen, DeepSeek, Kimi, and Llama variants rotate. Weights can be free. The hosted coding API still bills tokens.