Terminal-Bench 2.1

    🏆 Leaderboard

    As of August 20, 2026, GPT-5.6 Sol is #1 for Terminal-Bench 2.1 at 88.8%. Ranked by Terminal-Bench: can the model finish jobs in a real shell. 52 models in this index have a published Terminal-Bench 2.1 score. Methodology: Terminal-Bench (https://www.tbench.ai/). Terminal-Bench 2.1 leaderboard with live API prices. Terminal-Bench 2.1 — harder terminal and DevOps agent tasks than Terminal-Bench 2.0. Official methodology: Terminal-Bench (https://www.tbench.ai/).

    Updated August 20, 2026282 models33 providers
    GPT-5.6 Sol
    OpenAI · Proprietary
    88.8%$35.00
    Kimi K3OSS
    Moonshot AI · Open Source
    88.3%$18.00
    GLM-5.3OSS
    Z AI · Open Source
    88.2%$5.80
    Claude Mythos 5
    Anthropic · Proprietary
    88%$60.00
    DeepSeek-V4-Pro-0813OSS
    DeepSeek · Open Source
    87.9%$5.28
    GPT-5.6 Terra
    OpenAI · Proprietary
    87.4%$14.00
    Qwen3.8 MaxOSS
    Qwen · Open Source
    86.6%$8.00
    Gemini 3.7 Flash
    Google · Proprietary
    85.8%$4.50
    GPT-5.6 Luna
    OpenAI · Proprietary
    84.7%$1.40
    Claude Opus 5
    Anthropic · Proprietary
    84.6%$30.00
    Claude Fable 5
    Anthropic · Proprietary
    84.3%$60.00
    Grok 4.5
    xAI · Proprietary
    83.3%$8.00
    Muse Spark 1.2
    Meta · Proprietary
    82.9%$5.50
    GLM-5.2OSS
    Z AI · Open Source
    82.7%$5.80
    DeepSeek-V4-Flash-0731OSS
    DeepSeek · Open Source
    82.7%$1.76
    Muse Spark 1.1
    Meta · Proprietary
    80%$5.50
    Grok 4.6
    xAI · Proprietary
    78.3%$8.00
    Gemini 3.6 Flash
    Google · Proprietary
    78%$4.50
    GPT-5.5
    OpenAI · Proprietary
    76.4%$35.00
    Gemini 3.5 Flash
    Google · Proprietary
    76.2%$10.50
    Claude Sonnet 5
    Anthropic · Proprietary
    74.5%$12.00
    Qwen3.8-27BOSS
    Qwen · Open Source
    73%$3.65
    Claude Opus 4.8
    Anthropic · Proprietary
    71.9%$30.00
    Hy3OSS
    Tencent · Open Source
    71.7%$0.66
    Gemini 3.1 Pro
    Google · Proprietary
    70.8%$17.50
    Laguna S 2.1OSS
    Poolside · Open Source
    70.2%$0.30
    Claude Opus 4.7
    Anthropic · Proprietary
    68.5%$30.00
    Seed 2.1 Turbo
    ByteDance · Proprietary
    67.6%$3.00
    Kimi K2.7 CodeOSS
    Moonshot AI · Open Source
    67.0%$4.93
    MiniMax M3OSS
    MiniMax · Open Source
    66%$1.50
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    64.7%$1.50
    MAI-Code-1.1-Flash
    Microsoft · Proprietary
    62.9%$1.40
    Qwen3.7 Max
    Qwen · Proprietary
    61.0%$5.00
    MiMo-V2.5OSS
    Xiaomi · Open Source
    60.7%$0.50
    MiMo-V2.5-ProOSS
    Xiaomi · Open Source
    57.3%$1.30
    Claude Sonnet 4.6
    Anthropic · Proprietary
    57.3%$18.00
    Solar Pro 4
    Upstage · Proprietary
    57%$1.50
    GLM-5.1OSS
    Z AI · Open Source
    56.9%$5.80
    Nemotron 3 Ultra (550B A55B)OSS
    NVIDIA · Open Source · via OpenRouter
    56.4%$2.70
    GPT-5.4 mini
    OpenAI · Proprietary
    54.7%$5.25
    Gemini 3.5 Flash-Lite
    Google · Proprietary
    54%$2.80
    Gemini 3 Flash
    Google · Proprietary
    53.9%$3.50
    Kimi K2.6OSS
    Moonshot AI · Open Source
    53.6%$4.93
    Qwen3.6 Plus
    Qwen · Proprietary
    53.2%$3.50
    Qwen3.7-Plus
    Qwen · Proprietary
    52.8%$1.60
    Muse Glimmer-30BOSS
    Meta · Open Source
    51.7%$1.85
    MiniMax M2.7OSS
    MiniMax · Open Source
    48.7%$1.50
    Grok 4.3
    xAI · Proprietary
    42.0%$3.75
    GPT-5.4 nano
    OpenAI · Proprietary
    41.6%$1.45
    Mistral Medium 3.5OSS
    Mistral · Open Source · via Mistral AI
    39.0%$9.00
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads Terminal-Bench 2.1 right now?

    As of August 20, 2026, GPT-5.6 Sol by OpenAI is #1 for Terminal-Bench 2.1 at 88.8%. Ranked by Terminal-Bench: can the model finish jobs in a real shell. This board also tracks Terminal-Bench 2.1. Next on the same board: Kimi K3 and GLM-5.3. This terminal-bench 2.1 leaderboard ranks models by Terminal-Bench 2.1. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: Terminal-Bench (https://www.tbench.ai/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for terminal-bench 2.1. Ranked by Terminal-Bench: can the model finish jobs in a real shell. Input and output are dollars per million tokens.
    RankModelTerminal-Bench 2.1Input /MOutput /M
    1GPT-5.6 Sol88.8%$5.00$30.00
    2Kimi K388.3%$3.00$15.00
    3GLM-5.388.2%$1.40$4.40
    4Claude Mythos 588%$10.00$50.00
    5DeepSeek-V4-Pro-081387.9%$1.32$3.96
    6GPT-5.6 Terra87.4%$2.00$12.00
    7Qwen3.8 Max86.6%$2.00$6.00
    8Gemini 3.7 Flash85.8%$0.75$3.75

    Terminal-Bench 2.1 FAQ

    Who ranks #1 on the Terminal-Bench 2.1 leaderboard?

    As of August 20, 2026, GPT-5.6 Sol by OpenAI ranks #1 on Terminal-Bench 2.1 at 88.8%. API pricing is $5.00/M input and $30.00/M output.

    What are the top models on Terminal-Bench 2.1?

    The current Terminal-Bench 2.1 ranking as of August 20, 2026 is 1. GPT-5.6 Sol at 88.8%; 2. Kimi K3 at 88.3%; 3. GLM-5.3 at 88.2%.

    Which terminal-bench 2.1 model is the cheapest?

    Nemotron 3.5 Lightning (30B A3B) is the cheapest scored model on this terminal-bench 2.1 leaderboard at $0.05/M input and $0.20/M output ($0.25 blended). GPT-5.6 Sol still leads Terminal-Bench 2.1 at 88.8%.

    Should I always pick the #1 Terminal-Bench 2.1 model?

    Not automatically. GPT-5.6 Sol leads Terminal-Bench 2.1, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh Terminal-Bench 2.1 against input/output price, context window, and related evals.

    How often is the Terminal-Bench 2.1 leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is Terminal-Bench 2.1?

    Terminal-Bench 2.1 — harder terminal and DevOps agent tasks than Terminal-Bench 2.0. This page ranks models that have published a Terminal-Bench 2.1 score, with live API token prices on the same row. Official methodology: Terminal-Bench (https://www.tbench.ai/).

    Where is the Terminal-Bench 2.1 leaderboard?

    This page is the Terminal-Bench 2.1 leaderboard. Models are sorted by Terminal-Bench 2.1, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this Terminal-Bench 2.1 ranking different from the official board?

    The official Terminal-Bench page owns the methodology. This page keeps the published Terminal-Bench 2.1 score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://www.tbench.ai/