Who ranks #1 on the Open Source leaderboard?
As of August 20, 2026, Kimi K3 by Moonshot AI ranks #1 on GPQA at 93.5%. API pricing is $3.00/M input and $15.00/M output.
As of August 20, 2026, Kimi K3 is #1 for open-weight models at 93.5%. Ranked by GPQA: graduate-level science questions that resist simple search. Best open source LLM and best local LLM, ranked by GPQA, HLE, and SWE-bench. Compare Llama vs Qwen vs DeepSeek with hosted API prices.
Next step
Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.
The index
282 models across 33 providers. Search or jump to a lab — every model page stays linked here.







As of August 20, 2026, Kimi K3 by Moonshot AI is #1 for open-weight models at 93.5%. Ranked by GPQA: graduate-level science questions that resist simple search. This board also tracks GPQA, AIME 2025, SWE-bench Verified, LiveCodeBench. Next on the same board: Qwen3.8 Max and GLM-5.2. Related leaders: Kimi K2-Thinking-0905 on AIME 2025 at 100%; DeepSeek-V4-Pro-0813 on SWE-bench Verified at 96.4%. This open source leaderboard ranks models by GPQA, using open-weight models only. Scores come from public evals. Prices are the live API rates in the table above.
Sources: GPQA (GitHub) (https://github.com/idavidrein/gpqa); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)
| Rank | Model | GPQA | Input /M | Output /M |
|---|---|---|---|---|
| 1 | Kimi K3 | 93.5% | $3.00 | $15.00 |
| 2 | Qwen3.8 Max | 92.6% | $2.00 | $6.00 |
| 3 | GLM-5.2 | 91.2% | $1.40 | $4.40 |
| 4 | Kimi K2.6 | 90.5% | $0.96 | $3.97 |
| 5 | Hy3 | 90.4% | $0.13 | $0.53 |
| 6 | DeepSeek-V4-Pro-Max | 90.1% | $1.74 | $3.48 |
| 7 | Inkling-Small | 89.5% | $0.30 | $1.20 |
| 8 | Qwen3.8-27B | 89.2% | $0.45 | $3.20 |
Local means you run the weights. This page still shows hosted API prices because most teams try a host first. Rank GPQA, HLE, and SWE-bench among open weights, then check whether your GPU can actually load the checkpoint.
Qwen vs Llama is the comparison people type. DeepSeek and Kimi belong in the same list. Closed GPT and Claude rows are filtered out here on purpose. Sort the table. Brand threads on Reddit go stale in a week.
Rank one eval at a time. All LLM benchmarks.
As of August 20, 2026, Kimi K3 by Moonshot AI ranks #1 on GPQA at 93.5%. API pricing is $3.00/M input and $15.00/M output.
As of August 20, 2026, Kimi K3 by Moonshot AI is #1 for open-weight models at 93.5%. Ranked by GPQA: graduate-level science questions that resist simple search. This board also tracks GPQA, AIME 2025, SWE-bench Verified, LiveCodeBench. Next on the same board: Qwen3.8 Max and GLM-5.2. Related leaders: Kimi K2-Thinking-0905 on AIME 2025 at 100%; DeepSeek-V4-Pro-0813 on SWE-bench Verified at 96.4%.
The current GPQA ranking as of August 20, 2026 is 1. Kimi K3 at 93.5%; 2. Qwen3.8 Max at 92.6%; 3. GLM-5.2 at 91.2%.
Llama 3.2 3B Instruct is the cheapest scored model on this open source leaderboard at $0.01/M input and $0.02/M output ($0.03 blended). Kimi K3 still leads GPQA at 93.5%.
Not automatically. Kimi K3 leads GPQA, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh GPQA against input/output price, context window, and related evals.
Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.
As of August 20, 2026, Kimi K3 ranks first among open-weight models on GPQA at 93.5%. Hosted API pricing is $3.00/M input and $15.00/M output. Local means you run the weights. The API still bills tokens.
Weights can be free while the hosted API still bills tokens. This table uses the public API price we track, not your self-host electricity cost.
Yes. GPT, Claude, and Gemini are filtered out so the ranking is only open-weight models.
Llama vs Qwen depends on the eval and the size. Qwen often leads public reasoning and coding tables in this index. Llama still wins on distribution and tooling. Sort this open-weight board instead of picking a brand.
No. Open weights can be free to download. Hosted Llama, Qwen, DeepSeek, and Kimi APIs still bill tokens. This table uses the public hosted rate, not your electricity bill.
Yes. This ranking is open-weight only. Closed models stay on the overall and coding boards.
Sort this board by SWE-bench or open the coding board and filter to open weights. DeepSeek, Qwen, Kimi, and Llama rotate. Confirm the host in the row. Same weights, different APIs, different bills.