Terminal-Bench

    🏆 Leaderboard

    As of August 20, 2026, Kimi K3 is #1 for Terminal-Bench at 88.3%. Ranked by Terminal-Bench: can the model finish jobs in a real shell. 67 models in this index have a published Terminal-Bench score. Methodology: Terminal-Bench (https://www.tbench.ai/). Terminal-Bench leaderboard with live API prices. Terminal-Bench — tests an AI agent's ability to solve tasks using terminal commands and system administration. Official methodology: Terminal-Bench (https://www.tbench.ai/).

    Updated August 20, 2026282 models33 providers
    Kimi K3OSS
    Moonshot AI · Open Source
    88.3%$18.00
    GLM-5.3OSS
    Z AI · Open Source
    88.2%$5.80
    Claude Mythos 5
    Anthropic · Proprietary
    88%$60.00
    Qwen3.8 MaxOSS
    Qwen · Open Source
    86.6%$8.00
    GPT-5.5
    OpenAI · Proprietary
    84.7%$35.00
    Grok 4.5
    xAI · Proprietary
    83.3%$8.00
    GPT-5.4
    OpenAI · Proprietary
    81.8%$17.50
    Claude Sonnet 5
    Anthropic · Proprietary
    80.4%$12.00
    Gemini 3.1 Pro
    Google · Proprietary
    80.2%$17.50
    Claude Opus 4.7
    Anthropic · Proprietary
    80.2%$30.00
    Claude Opus 4.6
    Anthropic · Proprietary
    79.8%$30.00
    GPT-5.3 Codex
    OpenAI · Proprietary
    78.4%$15.75
    Gemini 3.6 Flash
    Google · Proprietary
    78%$4.50
    Gemini 3.5 Flash
    Google · Proprietary
    76.2%$10.50
    Qwen3.8-27BOSS
    Qwen · Open Source
    73%$3.65
    Gemini 3 Pro
    Google · Proprietary
    69.4%$14.00
    GPT-5.2 Codex
    OpenAI · Proprietary
    66.5%$15.75
    GPT-5.2
    OpenAI · Proprietary
    64.9%$15.75
    Gemini 3 Flash
    Google · Proprietary
    64.3%$3.50
    Claude Opus 4.5
    Anthropic · Proprietary
    63.1%$30.00
    GPT-5.1 Codex Mini
    OpenAI · Proprietary
    61.6%$2.25
    GPT-5.1 Codex
    OpenAI · Proprietary
    60.4%$11.25
    Gemini 3.5 Flash-Lite
    Google · Proprietary
    54%$2.80
    Claude Sonnet 4.6
    Anthropic · Proprietary
    53.4%$18.00
    GLM-5OSS
    Z AI · Open Source
    52.4%$4.20
    Claude Sonnet 4.5
    Anthropic · Proprietary
    50%$18.00
    GPT-5
    OpenAI · Proprietary
    49.6%$11.25
    MiniMax M2.1OSS
    MiniMax · Open Source
    47.9%$1.50
    GPT-5.1
    OpenAI · Proprietary
    47.6%$11.25
    Kimi K2-Thinking-0905OSS
    Moonshot AI · Open Source
    47.1%$2.47
    MiniMax M2OSS
    MiniMax · Open Source
    46.3%$1.50
    MiniMax M2.7OSS
    MiniMax · Open Source
    45.1%$1.50
    GPT-5 Codex
    OpenAI · Proprietary
    44.3%$11.25
    Claude Opus 4.1
    Anthropic · Proprietary
    43.3%$90.00
    Kimi K2.5OSS
    Moonshot AI · Open Source
    43.2%$3.68
    MiniMax M2.5OSS
    MiniMax · Open Source
    42.7%$1.50
    Claude Haiku 4.5
    Anthropic · Proprietary
    41%$6.00
    GLM-4.6OSS
    Z AI · Open Source
    40.5%$2.80
    DeepSeek-V3.2OSS
    DeepSeek · Open Source · via OpenRouter
    39.6%$0.57
    LongCat-Flash-ChatOSS
    Meituan · Open Source
    39.5%$1.50
    Claude Opus 4
    Anthropic · Proprietary
    39.2%$90.00
    DeepSeek-V3.2-ExpOSS
    DeepSeek · Open Source
    37.7%$0.68
    GLM-4.5OSS
    Z AI · Open Source
    37.5%$2.80
    Claude Sonnet 4
    Anthropic · Proprietary
    35.5%$18.00
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    35.2%$18.00
    GPT-5 mini
    OpenAI · Proprietary
    34.8%$2.25
    LongCat-Flash-LiteOSS
    Meituan · Open Source
    33.8%$0.50
    GLM-4.7OSS
    Z AI · Open Source
    33.3%$2.80
    Gemini 2.5 Pro
    Google · Proprietary
    32.6%$11.25
    Nova 2 Lite
    Amazon · Proprietary
    32.5%$2.80
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads Terminal-Bench right now?

    As of August 20, 2026, Kimi K3 by Moonshot AI is #1 for Terminal-Bench at 88.3%. Ranked by Terminal-Bench: can the model finish jobs in a real shell. This board also tracks Terminal-Bench. Next on the same board: GLM-5.3 and Claude Mythos 5. This terminal-bench leaderboard ranks models by Terminal-Bench. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: Terminal-Bench (https://www.tbench.ai/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for terminal-bench. Ranked by Terminal-Bench: can the model finish jobs in a real shell. Input and output are dollars per million tokens.
    RankModelTerminal-BenchInput /MOutput /M
    1Kimi K388.3%$3.00$15.00
    2GLM-5.388.2%$1.40$4.40
    3Claude Mythos 588%$10.00$50.00
    4Qwen3.8 Max86.6%$2.00$6.00
    5GPT-5.584.7%$5.00$30.00
    6Grok 4.583.3%$2.00$6.00
    7GPT-5.481.8%$2.50$15.00
    8Claude Sonnet 580.4%$2.00$10.00

    Terminal-Bench FAQ

    Who ranks #1 on the Terminal-Bench leaderboard?

    As of August 20, 2026, Kimi K3 by Moonshot AI ranks #1 on Terminal-Bench at 88.3%. API pricing is $3.00/M input and $15.00/M output.

    What are the top models on Terminal-Bench?

    The current Terminal-Bench ranking as of August 20, 2026 is 1. Kimi K3 at 88.3%; 2. GLM-5.3 at 88.2%; 3. Claude Mythos 5 at 88%.

    Which terminal-bench model is the cheapest?

    Qwen3.5-9B is the cheapest scored model on this terminal-bench leaderboard at $0.10/M input and $0.15/M output ($0.25 blended). Kimi K3 still leads Terminal-Bench at 88.3%.

    Should I always pick the #1 Terminal-Bench model?

    Not automatically. Kimi K3 leads Terminal-Bench, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh Terminal-Bench against input/output price, context window, and related evals.

    How often is the Terminal-Bench leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is Terminal-Bench?

    Terminal-Bench — tests an AI agent's ability to solve tasks using terminal commands and system administration. This page ranks models that have published a Terminal-Bench score, with live API token prices on the same row. Official methodology: Terminal-Bench (https://www.tbench.ai/).

    Where is the Terminal-Bench leaderboard?

    This page is the Terminal-Bench leaderboard. Models are sorted by Terminal-Bench, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this Terminal-Bench ranking different from the official board?

    The official Terminal-Bench page owns the methodology. This page keeps the published Terminal-Bench score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://www.tbench.ai/