Terminal-Bench

    🏆 Leaderboard

    As of August 20, 2026, GPT-5.5 is #1 for Terminal-Bench at 82.7%. Ranked by the Terminal-Bench 2.0 score Terminal-Bench leaderboard for 2.0 and 2.1 agent scores, plus API pricing. Rank the best terminal and DevOps LLMs.

    Updated August 20, 2026282 models33 providers
    GPT-5.5
    OpenAI · Proprietary
    82.7%84.7%76.4%82.7%$35.00
    Claude Mythos Preview
    Anthropic · Proprietary
    82%82%$60.00
    Claude Sonnet 5
    Anthropic · Proprietary
    80.4%80.4%74.5%80.4%$12.00
    Qwen3.7-Plus
    Qwen · Proprietary
    70.3%52.8%70.3%$1.60
    Claude Opus 4.8
    Anthropic · Proprietary
    70.0%71.9%74.6%$30.00
    Claude Opus 4.7
    Anthropic · Proprietary
    68.5%80.2%68.5%69.4%$30.00
    MiMo-V2.5-ProOSS
    Xiaomi · Open Source
    68.4%57.3%68.4%$1.30
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    67.9%67.9%$5.22
    Gemini 3.5 Flash
    Google · Proprietary
    67.4%76.2%76.2%76.2%$10.50
    Gemini 3.1 Pro
    Google · Proprietary
    67.4%80.2%70.8%68.5%$17.50
    MiMo-V2.5OSS
    Xiaomi · Open Source
    65.8%60.7%65.8%$0.50
    Claude Opus 4.6
    Anthropic · Proprietary
    65.4%79.8%65.4%$30.00
    GPT-5.2 Codex
    OpenAI · Proprietary
    64%66.5%64%$15.75
    GPT-5.4
    OpenAI · Proprietary
    62.2%81.8%75.1%$17.50
    Composer 2 Fast
    Cursor · Proprietary
    61.7%$9.00
    Composer 2
    Cursor · Proprietary
    61.7%$3.00
    GPT-5.4 mini
    OpenAI · Proprietary
    60%54.7%60%$5.25
    Claude Sonnet 4.6
    Anthropic · Proprietary
    59.5%53.4%57.3%59.1%$18.00
    Qwen3.7 Max
    Qwen · Proprietary
    59.2%61.0%69.7%$5.00
    Claude Opus 4.5
    Anthropic · Proprietary
    58.4%63.1%59.3%$30.00
    Kimi K2.6OSS
    Moonshot AI · Open Source
    57.3%53.6%66.7%$4.93
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    56.9%56.9%$0.42
    GPT-5.3 Codex
    OpenAI · Proprietary
    56.7%78.4%77.3%$15.75
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    56.6%56.6%$0.30
    GLM-5OSS
    Z AI · Open Source
    56.2%52.4%56.2%$4.20
    DeepSeek-V4-Pro-0813OSS
    DeepSeek · Open Source
    56.2%87.9%$5.28
    GLM-5.1OSS
    Z AI · Open Source
    53.9%56.9%69%$5.80
    GPT-5.1 Codex
    OpenAI · Proprietary
    52.8%60.4%52.8%$11.25
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    52.5%52.5%$4.20
    GPT-5.2
    OpenAI · Proprietary
    51.7%64.9%$15.75
    Gemini 3 Flash
    Google · Proprietary
    51.7%64.3%53.9%47.6%$3.50
    Qwen3.6-35B-A3BOSS
    Qwen · Open Source · via OpenRouter
    51.5%24.6%51.5%$1.14
    Step-3.5-FlashOSS
    StepFun · Open Source
    51%51%$0.50
    Kimi K2.5OSS
    Moonshot AI · Open Source
    50.8%43.2%50.8%$3.68
    Qwen3.5-122B-A10BOSS
    Qwen · Open Source
    49.4%49.4%$3.60
    MiniMax M2.7OSS
    MiniMax · Open Source
    47.2%45.1%48.7%57%$1.50
    DeepSeek-V3.2OSS
    DeepSeek · Open Source · via OpenRouter
    46.4%39.6%46.4%$0.57
    GPT-5.4 nano
    OpenAI · Proprietary
    46.3%41.6%46.3%$1.45
    MiniMax M3OSS
    MiniMax · Open Source
    46.1%66%$1.50
    Qwen3.6-27BOSS
    Qwen · Open Source
    44.9%59.3%$4.20
    Qwen3.6 Plus
    Qwen · Proprietary
    44.9%53.2%61.6%$3.50
    GPT-5.1
    OpenAI · Proprietary
    44.9%47.6%$11.25
    Grok 4.3
    xAI · Proprietary
    43.5%42.0%$3.75
    Qwen3.5-27BOSS
    Qwen · Open Source
    41.6%41.6%$2.70
    MiniMax M2.5OSS
    MiniMax · Open Source
    41.6%42.7%$1.50
    Qwen3.5-35B-A3BOSS
    Qwen · Open Source
    40.5%40.5%$2.25
    Gemma 4 31BOSS
    Google · Open Source
    39.3%$0.54
    MiMo-V2-FlashOSS
    Xiaomi · Open Source
    38.5%30.5%38.5%$0.40
    GLM-4.7OSS
    Z AI · Open Source
    38.2%33.3%41%$2.80
    Qwen3-Coder 480B A35B InstructOSS
    Qwen · Open Source · via OpenRouter
    37.5%27.2%37.5%$2.02
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads Terminal-Bench right now?

    As of August 20, 2026, GPT-5.5 by OpenAI is #1 for Terminal-Bench at 82.7%. Ranked by the Terminal-Bench 2.0 score This board also tracks Terminal-Bench 2.0, Terminal-Bench, Terminal-Bench 2.1, Terminal-Bench 2. Next on the same board: Claude Mythos Preview and Claude Sonnet 5. Related leaders: Kimi K3 on Terminal-Bench at 88.3%; GPT-5.6 Sol on Terminal-Bench 2.1 at 88.8%. This terminal-bench leaderboard ranks models by Terminal-Bench 2.0. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: Terminal-Bench (https://www.tbench.ai/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for terminal-bench. Ranked by the Terminal-Bench 2.0 score Input and output are dollars per million tokens.
    RankModelTerminal-Bench 2.0Input /MOutput /M
    1GPT-5.582.7%$5.00$30.00
    2Claude Mythos Preview82%$10.00$50.00
    3Claude Sonnet 580.4%$2.00$10.00
    4Qwen3.7-Plus70.3%$0.32$1.28
    5Claude Opus 4.870.0%$5.00$25.00
    6Claude Opus 4.768.5%$5.00$25.00
    7MiMo-V2.5-Pro68.4%$0.43$0.87
    8DeepSeek-V4-Pro-Max67.9%$1.74$3.48

    Terminal-Bench FAQ

    Who ranks #1 on the Terminal-Bench leaderboard?

    As of August 20, 2026, GPT-5.5 by OpenAI ranks #1 on Terminal-Bench 2.0 at 82.7%. API pricing is $5.00/M input and $30.00/M output.

    What are the top models on Terminal-Bench 2.0?

    The current Terminal-Bench 2.0 ranking as of August 20, 2026 is 1. GPT-5.5 at 82.7%; 2. Claude Mythos Preview at 82%; 3. Claude Sonnet 5 at 80.4%.

    Which terminal-bench model is the cheapest?

    DeepSeek-V4-Flash-0423 is the cheapest scored model on this terminal-bench leaderboard at $0.10/M input and $0.20/M output ($0.30 blended). GPT-5.5 still leads Terminal-Bench 2.0 at 82.7%.

    Should I always pick the #1 Terminal-Bench 2.0 model?

    Not automatically. GPT-5.5 leads Terminal-Bench 2.0, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh Terminal-Bench 2.0 against input/output price, context window, and related evals.

    How often is the Terminal-Bench leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is Terminal-Bench?

    Terminal-Bench measures whether an agent can finish realistic shell, sysadmin, and DevOps tasks in a terminal instead of writing a single function.

    Should I use Terminal-Bench 2.0 or 2.1?

    Use 2.1 when the model has a score. It is the harder current set. We still show 2.0 and the original Terminal-Bench when that is all a lab published.

    Is Terminal-Bench the same as SWE-bench?

    No. SWE-bench is repository issue fixing. Terminal-Bench is command-line operations. Strong coding models can still fail basic shell workflows.