Open Source

    🏆 Leaderboard

    As of August 20, 2026, Kimi K3 is #1 for open-weight models at 93.5%. Ranked by GPQA: graduate-level science questions that resist simple search. Best open source LLM and best local LLM, ranked by GPQA, HLE, and SWE-bench. Compare Llama vs Qwen vs DeepSeek with hosted API prices.

    Updated August 20, 2026282 models33 providers
    Kimi K3OSS
    Moonshot AI · Open Source
    93.5%68.9%93.4%87.2%56%$18.00
    Qwen3.8 MaxOSS
    Qwen · Open Source
    92.6%99.4%85.6%87.8%43.6%$8.00
    GLM-5.2OSS
    Z AI · Open Source
    91.2%28.9%78.7%69.5%54.7%$5.80
    Kimi K2.6OSS
    Moonshot AI · Open Source
    90.5%96.1%80.2%86.8%36.4%$4.93
    Hy3OSS
    Tencent · Open Source
    90.4%78%$0.66
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    90.1%96.7%80.6%93.5%48.2%$5.22
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    89.5%90%80.2%85.9%31.6%$1.50
    Qwen3.8-27BOSS
    Qwen · Open Source
    89.2%30.8%$3.65
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    88.4%88.9%76.4%28.7%87.8%$4.20
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    88.1%79%91.6%45.1%$0.42
    Qwen3.6-27BOSS
    Qwen · Open Source
    87.8%91.1%77.2%24%$4.20
    Kimi K2.5OSS
    Moonshot AI · Open Source
    87.6%96.1%76.8%50.2%$3.68
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    87.4%78.6%88.4%40.3%$0.30
    Nemotron 3 Ultra (550B A55B)OSS
    NVIDIA · Open Source · via OpenRouter
    87%70.7%86.0%37.4%$2.70
    Qwen3.5-122B-A10BOSS
    Qwen · Open Source
    86.6%72%47.5%$3.60
    GLM-5.1OSS
    Z AI · Open Source
    86.2%93.3%74.2%81.4%52.3%$5.80
    Qwen3.6-35B-A3BOSS
    Qwen · Open Source · via OpenRouter
    86%86.7%73.4%21.4%$1.14
    GLM-4.7OSS
    Z AI · Open Source
    85.7%95.7%73.8%82.2%42.8%84.3%$2.80
    Qwen3.5-27BOSS
    Qwen · Open Source
    85.5%72.4%48.5%$2.70
    Kimi K2-Thinking-0905OSS
    Moonshot AI · Open Source
    84.5%100%71.3%51%$2.47
    Gemma 4 31BOSS
    Google · Open Source
    84.3%26.5%$0.54
    Qwen3.5-35B-A3BOSS
    Qwen · Open Source
    84.2%69.2%47.4%$2.25
    MiMo-V2-FlashOSS
    Xiaomi · Open Source
    83.7%94.1%73.4%22.1%$0.40
    Muse Glimmer-30BOSS
    Meta · Open Source
    83.5%76%22%$1.85
    Nemotron 3 Super (120B A12B)OSS
    NVIDIA · Open Source · via OpenRouter
    82.7%90.2%53.7%81.2%22.8%$0.54
    DeepSeek-V3.2OSS
    DeepSeek · Open Source · via OpenRouter
    82.4%93.1%73.1%83.3%40.8%$0.57
    Gemma 4 26B-A4BOSS
    Google · Open Source
    82.3%17.2%$0.53
    Qwen3.5-9BOSS
    Qwen · Open Source · via OpenRouter
    81.7%$0.25
    LongCat-Flash-ThinkingOSS
    Meituan · Open Source
    81.5%90.6%59.4%79.4%$1.50
    Qwen3-235B-A22B-Thinking-2507OSS
    Qwen · Open Source
    81.1%92.3%18.2%$3.30
    MiniMax M2.1OSS
    MiniMax · Open Source
    81%81%67%78%22%$1.50
    GLM-4.6OSS
    Z AI · Open Source
    81%93.9%68%81.0%17.2%$2.80
    DeepSeek-R1-0528OSS
    DeepSeek · Open Source
    81%87.5%44.6%73.3%17.7%$2.74
    GPT OSS 120B HighOSS
    OpenAI · Open Source
    80.9%92.5%$0.60
    LongCat-Flash-Thinking-2601OSS
    Meituan · Open Source
    80.5%99.6%70%82.8%25.2%$1.50
    GPT OSS 120BOSS
    OpenAI · Open Source
    80.1%92.6%33.6%83.2%14.9%90%$0.54
    DeepSeek-V3.2-ExpOSS
    DeepSeek · Open Source
    79.9%89.3%67.8%74.1%19.8%$0.68
    GLM-4.5OSS
    Z AI · Open Source
    79.1%86.7%64.2%72.9%14.4%84.6%$2.80
    MiniMax M2OSS
    MiniMax · Open Source
    78%78%69.4%83%12.5%$1.50
    Qwen3-235B-A22B-Instruct-2507OSS
    Qwen · Open Source
    77.5%70.3%$0.95
    Qwen3-Next-80B-A3B-ThinkingOSS
    Qwen · Open Source
    77.2%87.8%$1.65
    Nemotron 3.5 Lightning (30B A3B)OSS
    NVIDIA · Open Source
    75.4%51.6%11.7%$0.25
    GLM-4.7-FlashOSS
    Z AI · Open Source
    75.2%91.6%59.2%14.4%
    Kimi K2-Instruct-0905OSS
    Moonshot AI · Open Source · via OpenRouter
    75.1%49.5%65.8%53.7%4.7%89.5%$3.10
    Kimi K2 InstructOSS
    Moonshot AI · Open Source
    75.1%49.5%43.8%70.5%4.7%89.5%$1.00
    Nemotron 3 Nano (30B A3B)OSS
    NVIDIA · Open Source
    75%99.2%38.8%15.5%$0.30
    GLM-4.5-AirOSS
    Z AI · Open Source
    75%57.6%70.7%10.6%$1.30
    DeepSeek-V3.1OSS
    DeepSeek · Open Source
    74.9%49.8%66%56.4%15.9%$1.27
    Qwen3 VL 30B A3B ThinkingOSS
    Qwen · Open Source
    74.4%83.1%87.6%$1.20
    GPT OSS 20B HighOSS
    OpenAI · Open Source
    74.2%98.7%$0.60
    Showing 150 of 154 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    What is the best open source LLM?

    As of August 20, 2026, Kimi K3 by Moonshot AI is #1 for open-weight models at 93.5%. Ranked by GPQA: graduate-level science questions that resist simple search. This board also tracks GPQA, AIME 2025, SWE-bench Verified, LiveCodeBench. Next on the same board: Qwen3.8 Max and GLM-5.2. Related leaders: Kimi K2-Thinking-0905 on AIME 2025 at 100%; DeepSeek-V4-Pro-0813 on SWE-bench Verified at 96.4%. This open source leaderboard ranks models by GPQA, using open-weight models only. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: GPQA (GitHub) (https://github.com/idavidrein/gpqa); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for open source. Ranked by GPQA: graduate-level science questions that resist simple search. Input and output are dollars per million tokens.
    RankModelGPQAInput /MOutput /M
    1Kimi K393.5%$3.00$15.00
    2Qwen3.8 Max92.6%$2.00$6.00
    3GLM-5.291.2%$1.40$4.40
    4Kimi K2.690.5%$0.96$3.97
    5Hy390.4%$0.13$0.53
    6DeepSeek-V4-Pro-Max90.1%$1.74$3.48
    7Inkling-Small89.5%$0.30$1.20
    8Qwen3.8-27B89.2%$0.45$3.20

    What is the best local LLM?

    Local means you run the weights. This page still shows hosted API prices because most teams try a host first. Rank GPQA, HLE, and SWE-bench among open weights, then check whether your GPU can actually load the checkpoint.

    Llama vs Qwen vs DeepSeek

    Qwen vs Llama is the comparison people type. DeepSeek and Kimi belong in the same list. Closed GPT and Claude rows are filtered out here on purpose. Sort the table. Brand threads on Reddit go stale in a week.

    Open Source FAQ

    Who ranks #1 on the Open Source leaderboard?

    As of August 20, 2026, Kimi K3 by Moonshot AI ranks #1 on GPQA at 93.5%. API pricing is $3.00/M input and $15.00/M output.

    What is the best open source LLM?

    As of August 20, 2026, Kimi K3 by Moonshot AI is #1 for open-weight models at 93.5%. Ranked by GPQA: graduate-level science questions that resist simple search. This board also tracks GPQA, AIME 2025, SWE-bench Verified, LiveCodeBench. Next on the same board: Qwen3.8 Max and GLM-5.2. Related leaders: Kimi K2-Thinking-0905 on AIME 2025 at 100%; DeepSeek-V4-Pro-0813 on SWE-bench Verified at 96.4%.

    What are the top models on GPQA?

    The current GPQA ranking as of August 20, 2026 is 1. Kimi K3 at 93.5%; 2. Qwen3.8 Max at 92.6%; 3. GLM-5.2 at 91.2%.

    Which open source model is the cheapest?

    Llama 3.2 3B Instruct is the cheapest scored model on this open source leaderboard at $0.01/M input and $0.02/M output ($0.03 blended). Kimi K3 still leads GPQA at 93.5%.

    Should I always pick the #1 GPQA model?

    Not automatically. Kimi K3 leads GPQA, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh GPQA against input/output price, context window, and related evals.

    How often is the Open Source leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is the best local LLM?

    As of August 20, 2026, Kimi K3 ranks first among open-weight models on GPQA at 93.5%. Hosted API pricing is $3.00/M input and $15.00/M output. Local means you run the weights. The API still bills tokens.

    Does open source mean free to run?

    Weights can be free while the hosted API still bills tokens. This table uses the public API price we track, not your self-host electricity cost.

    Are closed models hidden on this page?

    Yes. GPT, Claude, and Gemini are filtered out so the ranking is only open-weight models.

    Llama vs Qwen: which open source model is better?

    Llama vs Qwen depends on the eval and the size. Qwen often leads public reasoning and coding tables in this index. Llama still wins on distribution and tooling. Sort this open-weight board instead of picking a brand.

    Does open source mean the API is free?

    No. Open weights can be free to download. Hosted Llama, Qwen, DeepSeek, and Kimi APIs still bill tokens. This table uses the public hosted rate, not your electricity bill.

    Are GPT and Claude hidden here?

    Yes. This ranking is open-weight only. Closed models stay on the overall and coding boards.

    What is the best open source coding LLM?

    Sort this board by SWE-bench or open the coding board and filter to open weights. DeepSeek, Qwen, Kimi, and Llama rotate. Confirm the host in the row. Same weights, different APIs, different bills.