tau-bench Retail

    🏆 Leaderboard

    As of August 20, 2026, Claude Opus 4.6 is #1 for tau-bench Retail at 91.9%. Ranked by tau-bench: multi-turn tool use on customer-support workflows. 44 models in this index have a published tau-bench Retail score. Methodology: τ-bench (https://github.com/sierra-research/tau-bench). tau-bench Retail leaderboard with live API prices. tau-bench Retail — tests AI agents on realistic retail customer service interactions. Official methodology: τ-bench (https://github.com/sierra-research/tau-bench).

    Updated August 20, 2026282 models33 providers
    Claude Opus 4.6
    Anthropic · Proprietary
    91.9%$30.00
    Claude Sonnet 4.6
    Anthropic · Proprietary
    91.7%$18.00
    Claude Opus 4.5
    Anthropic · Proprietary
    88.9%$30.00
    LongCat-Flash-Thinking-2601OSS
    Meituan · Open Source
    88.6%$1.50
    Claude Sonnet 4.5
    Anthropic · Proprietary
    86.2%$18.00
    Claude Haiku 4.5
    Anthropic · Proprietary
    83.2%$6.00
    Claude Opus 4.1
    Anthropic · Proprietary
    82.4%$90.00
    GPT-5.2
    OpenAI · Proprietary
    82%$15.75
    Claude Opus 4
    Anthropic · Proprietary
    81.4%$90.00
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    81.2%$18.00
    GPT-5
    OpenAI · Proprietary
    81.1%$11.25
    Claude Sonnet 4
    Anthropic · Proprietary
    80.5%$18.00
    o3
    OpenAI · Proprietary
    80.2%$10.00
    GLM-4.5OSS
    Z AI · Open Source
    79.7%$2.80
    GPT-5.1 Thinking
    OpenAI · Proprietary
    77.9%$11.25
    GPT-5.1 Instant
    OpenAI · Proprietary
    77.9%$11.25
    GPT-5.1
    OpenAI · Proprietary
    77.9%$11.25
    GLM-4.5-AirOSS
    Z AI · Open Source
    77.9%$1.30
    Qwen3-Coder 480B A35B InstructOSS
    Qwen · Open Source · via OpenRouter
    77.5%$2.02
    Nova 2 Lite
    Amazon · Proprietary
    76.5%$2.80
    LongCat-Flash-LiteOSS
    Meituan · Open Source
    73.1%$0.50
    o4-mini
    OpenAI · Proprietary
    71.8%$5.50
    LongCat-Flash-ThinkingOSS
    Meituan · Open Source
    71.5%$1.50
    Qwen3-235B-A22B-Instruct-2507OSS
    Qwen · Open Source
    71.3%$0.95
    LongCat-Flash-ChatOSS
    Meituan · Open Source
    71.3%$1.50
    o1
    OpenAI · Proprietary
    70.8%$75.00
    Kimi K2-Instruct-0905OSS
    Moonshot AI · Open Source · via OpenRouter
    70.6%$3.10
    Kimi K2 InstructOSS
    Moonshot AI · Open Source
    70.6%$1.00
    Qwen3-Next-80B-A3B-ThinkingOSS
    Qwen · Open Source
    69.6%$1.65
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    69.2%$18.00
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    69.2%$18.00
    GPT-4.1
    OpenAI · Proprietary
    68%$10.00
    Qwen3-235B-A22B-Thinking-2507OSS
    Qwen · Open Source
    67.8%$3.30
    GPT OSS 120BOSS
    OpenAI · Open Source
    67.8%$0.54
    MiniMax M1 80KOSS
    MiniMax · Open Source
    63.5%$2.75
    Nemotron 3 Super (120B A12B)OSS
    NVIDIA · Open Source · via OpenRouter
    62.8%$0.54
    Qwen3-Next-80B-A3B-InstructOSS
    Qwen · Open Source
    60.9%$1.65
    GPT-4o
    OpenAI · Proprietary
    60.3%$12.50
    o3-mini
    OpenAI · Proprietary
    57.6%$5.50
    Nemotron 3 Nano (30B A3B)OSS
    NVIDIA · Open Source
    56.9%$0.30
    GPT-4.1 mini
    OpenAI · Proprietary
    55.8%$2.00
    GPT OSS 20BOSS
    OpenAI · Open Source
    54.8%$0.25
    Claude 3.5 Haiku
    Anthropic · Proprietary
    51%$4.80
    GPT-4.1 nano
    OpenAI · Proprietary
    22.6%$0.50
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Claude 3 Haiku
    Anthropic · Proprietary
    $1.50
    Claude 3 Opus
    Anthropic · Proprietary
    $90.00
    Claude 3 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude Fable 5
    Anthropic · Proprietary
    $60.00
    Claude Mythos 5
    Anthropic · Proprietary
    $60.00
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads tau-bench Retail right now?

    As of August 20, 2026, Claude Opus 4.6 by Anthropic is #1 for tau-bench Retail at 91.9%. Ranked by tau-bench: multi-turn tool use on customer-support workflows. This board also tracks tau-bench Retail. Next on the same board: Claude Sonnet 4.6 and Claude Opus 4.5. This tau-bench retail leaderboard ranks models by tau-bench Retail. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: τ-bench (https://github.com/sierra-research/tau-bench); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for tau-bench retail. Ranked by tau-bench: multi-turn tool use on customer-support workflows. Input and output are dollars per million tokens.
    RankModeltau-bench RetailInput /MOutput /M
    1Claude Opus 4.691.9%$5.00$25.00
    2Claude Sonnet 4.691.7%$3.00$15.00
    3Claude Opus 4.588.9%$5.00$25.00
    4LongCat-Flash-Thinking-260188.6%$0.30$1.20
    5Claude Sonnet 4.586.2%$3.00$15.00
    6Claude Haiku 4.583.2%$1.00$5.00
    7Claude Opus 4.182.4%$15.00$75.00
    8GPT-5.282%$1.75$14.00

    tau-bench Retail FAQ

    Who ranks #1 on the tau-bench Retail leaderboard?

    As of August 20, 2026, Claude Opus 4.6 by Anthropic ranks #1 on tau-bench Retail at 91.9%. API pricing is $5.00/M input and $25.00/M output.

    What are the top models on tau-bench Retail?

    The current tau-bench Retail ranking as of August 20, 2026 is 1. Claude Opus 4.6 at 91.9%; 2. Claude Sonnet 4.6 at 91.7%; 3. Claude Opus 4.5 at 88.9%.

    Which tau-bench retail model is the cheapest?

    GPT OSS 20B is the cheapest scored model on this tau-bench retail leaderboard at $0.05/M input and $0.20/M output ($0.25 blended). Claude Opus 4.6 still leads tau-bench Retail at 91.9%.

    Should I always pick the #1 tau-bench Retail model?

    Not automatically. Claude Opus 4.6 leads tau-bench Retail, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh tau-bench Retail against input/output price, context window, and related evals.

    How often is the tau-bench Retail leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is tau-bench Retail?

    tau-bench Retail — tests AI agents on realistic retail customer service interactions. This page ranks models that have published a tau-bench Retail score, with live API token prices on the same row. Official methodology: τ-bench (https://github.com/sierra-research/tau-bench).

    Where is the tau-bench Retail leaderboard?

    This page is the tau-bench Retail leaderboard. Models are sorted by tau-bench Retail, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this tau-bench Retail ranking different from the official board?

    The official τ-bench page owns the methodology. This page keeps the published tau-bench Retail score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://github.com/sierra-research/tau-bench