SWE-bench Multilingual

    🏆 Leaderboard

    As of August 20, 2026, Claude Mythos Preview is #1 for SWE-bench Multilingual at 87.3%. Ranked by the SWE-bench Multilingual score 38 models in this index have a published SWE-bench Multilingual score. SWE-bench Multilingual leaderboard with live API prices. SWE-bench Multilingual — a multilingual software engineering benchmark derived from SWE-bench tasks across multiple languages and locales.

    Updated August 20, 2026282 models33 providers
    Claude Mythos Preview
    Anthropic · Proprietary
    87.3%$60.00
    Claude Opus 4.8
    Anthropic · Proprietary
    84.4%$30.00
    Laguna S 2.1OSS
    Poolside · Open Source
    78.5%$0.30
    Qwen3.7 Max
    Qwen · Proprietary
    78.3%$5.00
    Claude Sonnet 5
    Anthropic · Proprietary
    78.3%$12.00
    Claude Opus 4.6
    Anthropic · Proprietary
    77.8%$30.00
    Claude Mythos 5
    Anthropic · Proprietary
    77.8%$60.00
    Kimi K2.6OSS
    Moonshot AI · Open Source
    76.7%$4.93
    MiniMax M2.7OSS
    MiniMax · Open Source
    76.5%$1.50
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    76.2%$5.22
    Qwen3.7-Plus
    Qwen · Proprietary
    75.8%$1.60
    Hy3OSS
    Tencent · Open Source
    75.8%$0.66
    Qwen3.6 Plus
    Qwen · Proprietary
    73.8%$3.50
    Composer 2 Fast
    Cursor · Proprietary
    73.7%$9.00
    Composer 2
    Cursor · Proprietary
    73.7%$3.00
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    73.3%$0.42
    Kimi K2.5OSS
    Moonshot AI · Open Source
    73%$3.68
    MiniMax M2.1OSS
    MiniMax · Open Source
    72.5%$1.50
    MiMo-V2-FlashOSS
    Xiaomi · Open Source
    71.7%$0.40
    Qwen3.6-27BOSS
    Qwen · Open Source
    71.3%$4.20
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    70.2%$0.30
    DeepSeek-V3.2OSS
    DeepSeek · Open Source · via OpenRouter
    70.2%$0.57
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    69.3%$4.20
    Nemotron 3 Ultra (550B A55B)OSS
    NVIDIA · Open Source · via OpenRouter
    67.7%$2.70
    Qwen3.6-35B-A3BOSS
    Qwen · Open Source · via OpenRouter
    67.2%$1.14
    GLM-4.7OSS
    Z AI · Open Source
    66.7%$2.80
    Laguna XS 2.1OSS
    Poolside · Open Source
    63.1%$0.30
    Kimi K2-Thinking-0905OSS
    Moonshot AI · Open Source
    61.1%$2.47
    DeepSeek-V3.2-ExpOSS
    DeepSeek · Open Source
    57.9%$0.68
    MiniMax M2OSS
    MiniMax · Open Source
    56.5%$1.50
    Qwen3-Coder 480B A35B InstructOSS
    Qwen · Open Source · via OpenRouter
    54.7%$2.02
    DeepSeek-V3.1OSS
    DeepSeek · Open Source
    54.5%$1.27
    Kimi K2-Instruct-0905OSS
    Moonshot AI · Open Source · via OpenRouter
    47.3%$3.10
    Kimi K2 InstructOSS
    Moonshot AI · Open Source
    47.3%$1.00
    Nemotron 3 Super (120B A12B)OSS
    NVIDIA · Open Source · via OpenRouter
    45.8%$0.54
    Nemotron 3.5 Lightning (30B A3B)OSS
    NVIDIA · Open Source
    39.3%$0.25
    LongCat-Flash-LiteOSS
    Meituan · Open Source
    38.1%$0.50
    DeepSeek-R1-0528OSS
    DeepSeek · Open Source
    30.5%$2.74
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Claude 3 Haiku
    Anthropic · Proprietary
    $1.50
    Claude 3 Opus
    Anthropic · Proprietary
    $90.00
    Claude 3 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Haiku
    Anthropic · Proprietary
    $4.80
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude Fable 5
    Anthropic · Proprietary
    $60.00
    Claude Haiku 4.5
    Anthropic · Proprietary
    $6.00
    Claude Opus 4
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.1
    Anthropic · Proprietary
    $90.00
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads SWE-bench Multilingual right now?

    As of August 20, 2026, Claude Mythos Preview by Anthropic is #1 for SWE-bench Multilingual at 87.3%. Ranked by the SWE-bench Multilingual score This board also tracks SWE-bench Multilingual. Next on the same board: Claude Opus 4.8 and Laguna S 2.1. This swe-bench multilingual leaderboard ranks models by SWE-bench Multilingual. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for swe-bench multilingual. Ranked by the SWE-bench Multilingual score Input and output are dollars per million tokens.
    RankModelSWE-bench MultilingualInput /MOutput /M
    1Claude Mythos Preview87.3%$10.00$50.00
    2Claude Opus 4.884.4%$5.00$25.00
    3Laguna S 2.178.5%$0.10$0.20
    4Claude Sonnet 578.3%$2.00$10.00
    5Qwen3.7 Max78.3%$1.25$3.75
    6Claude Opus 4.677.8%$5.00$25.00
    7Claude Mythos 577.8%$10.00$50.00
    8Kimi K2.676.7%$0.96$3.97

    SWE-bench Multilingual FAQ

    Who ranks #1 on the SWE-bench Multilingual leaderboard?

    As of August 20, 2026, Claude Mythos Preview by Anthropic ranks #1 on SWE-bench Multilingual at 87.3%. API pricing is $10.00/M input and $50.00/M output.

    What are the top models on SWE-bench Multilingual?

    The current SWE-bench Multilingual ranking as of August 20, 2026 is 1. Claude Mythos Preview at 87.3%; 2. Claude Opus 4.8 at 84.4%; 3. Laguna S 2.1 at 78.5%.

    Which swe-bench multilingual model is the cheapest?

    Nemotron 3.5 Lightning (30B A3B) is the cheapest scored model on this swe-bench multilingual leaderboard at $0.05/M input and $0.20/M output ($0.25 blended). Claude Mythos Preview still leads SWE-bench Multilingual at 87.3%.

    Should I always pick the #1 SWE-bench Multilingual model?

    Not automatically. Claude Mythos Preview leads SWE-bench Multilingual, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh SWE-bench Multilingual against input/output price, context window, and related evals.

    How often is the SWE-bench Multilingual leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is SWE-bench Multilingual?

    SWE-bench Multilingual — a multilingual software engineering benchmark derived from SWE-bench tasks across multiple languages and locales. This page ranks models that have published a SWE-bench Multilingual score, with live API token prices on the same row.

    Where is the SWE-bench Multilingual leaderboard?

    This page is the SWE-bench Multilingual leaderboard. Models are sorted by SWE-bench Multilingual, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this SWE-bench Multilingual ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published SWE-bench Multilingual score next to live API $/M so you can pick a production SKU, not only a trophy number.