tau2-bench Airline

    🏆 Leaderboard

    As of August 20, 2026, LongCat-Flash-Thinking-2601 is #1 for tau2-bench Airline at 76.5%. Ranked by the tau2-bench Airline score 21 models in this index have a published tau2-bench Airline score. tau2-bench Airline leaderboard: rank models by tau2-bench Airline next to live API token prices.

    Updated August 20, 2026282 models33 providers
    LongCat-Flash-Thinking-2601OSS
    Meituan · Open Source
    76.5%$1.50
    LongCat-Flash-ThinkingOSS
    Meituan · Open Source
    67.5%$1.50
    GPT-5.1 Thinking
    OpenAI · Proprietary
    67%$11.25
    GPT-5.1 Instant
    OpenAI · Proprietary
    67%$11.25
    GPT-5.1
    OpenAI · Proprietary
    67%$11.25
    o3
    OpenAI · Proprietary
    64.8%$10.00
    Nova 2 Lite
    Amazon · Proprietary
    64.8%$2.80
    Claude Haiku 4.5
    Anthropic · Proprietary
    63.6%$6.00
    GPT-5
    OpenAI · Proprietary
    62.6%$11.25
    Qwen3-Next-80B-A3B-ThinkingOSS
    Qwen · Open Source
    60.5%$1.65
    Qwen3-235B-A22B-Thinking-2507OSS
    Qwen · Open Source
    58%$3.30
    LongCat-Flash-LiteOSS
    Meituan · Open Source
    58%$0.50
    LongCat-Flash-ChatOSS
    Meituan · Open Source
    58%$1.50
    Kimi K2-Instruct-0905OSS
    Moonshot AI · Open Source · via OpenRouter
    56.5%$3.10
    Kimi K2 InstructOSS
    Moonshot AI · Open Source
    56.5%$1.00
    Nemotron 3 Super (120B A12B)OSS
    NVIDIA · Open Source · via OpenRouter
    56.3%$0.54
    Mercury 2
    Inception · Proprietary
    53%$1.00
    Nemotron 3 Nano (30B A3B)OSS
    NVIDIA · Open Source
    48%$0.30
    Qwen3-Next-80B-A3B-InstructOSS
    Qwen · Open Source
    45.5%$1.65
    GPT-4o
    OpenAI · Proprietary
    45.5%$12.50
    Qwen3-235B-A22B-Instruct-2507OSS
    Qwen · Open Source
    44%$0.95
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Claude 3 Haiku
    Anthropic · Proprietary
    $1.50
    Claude 3 Opus
    Anthropic · Proprietary
    $90.00
    Claude 3 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Haiku
    Anthropic · Proprietary
    $4.80
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude Fable 5
    Anthropic · Proprietary
    $60.00
    Claude Mythos 5
    Anthropic · Proprietary
    $60.00
    Claude Mythos Preview
    Anthropic · Proprietary
    $60.00
    Claude Opus 4
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.1
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.5
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.6
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.7
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.8
    Anthropic · Proprietary
    $30.00
    Claude Opus 5
    Anthropic · Proprietary
    $30.00
    Claude Sonnet 4
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 4.5
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 4.6
    Anthropic · Proprietary
    $18.00
    Claude Sonnet 5
    Anthropic · Proprietary
    $12.00
    Command A+OSS
    Cohere · Open Source
    $12.50
    Command R+OSS
    Cohere · Open Source
    $1.25
    Composer 2
    Cursor · Proprietary
    $3.00
    Composer 2 Fast
    Cursor · Proprietary
    $9.00
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek · Open Source
    $0.50
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek · Open Source
    $0.30
    DeepSeek-R1OSS
    DeepSeek · Open Source
    $2.74
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads tau2-bench Airline right now?

    As of August 20, 2026, LongCat-Flash-Thinking-2601 by Meituan is #1 for tau2-bench Airline at 76.5%. Ranked by the tau2-bench Airline score This board also tracks tau2-bench Airline. Next on the same board: LongCat-Flash-Thinking and GPT-5.1. This tau2-bench airline leaderboard ranks models by tau2-bench Airline. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for tau2-bench airline. Ranked by the tau2-bench Airline score Input and output are dollars per million tokens.
    RankModeltau2-bench AirlineInput /MOutput /M
    1LongCat-Flash-Thinking-260176.5%$0.30$1.20
    2LongCat-Flash-Thinking67.5%$0.30$1.20
    3GPT-5.167%$1.25$10.00
    4GPT-5.1 Instant67%$1.25$10.00
    5GPT-5.1 Thinking67%$1.25$10.00
    6Nova 2 Lite64.8%$0.30$2.50
    7o364.8%$2.00$8.00
    8Claude Haiku 4.563.6%$1.00$5.00

    tau2-bench Airline FAQ

    Who ranks #1 on the tau2-bench Airline leaderboard?

    As of August 20, 2026, LongCat-Flash-Thinking-2601 by Meituan ranks #1 on tau2-bench Airline at 76.5%. API pricing is $0.30/M input and $1.20/M output.

    What are the top models on tau2-bench Airline?

    The current tau2-bench Airline ranking as of August 20, 2026 is 1. LongCat-Flash-Thinking-2601 at 76.5%; 2. LongCat-Flash-Thinking at 67.5%; 3. GPT-5.1 at 67%.

    Which tau2-bench airline model is the cheapest?

    Nemotron 3 Nano (30B A3B) is the cheapest scored model on this tau2-bench airline leaderboard at $0.06/M input and $0.24/M output ($0.30 blended). LongCat-Flash-Thinking-2601 still leads tau2-bench Airline at 76.5%.

    Should I always pick the #1 tau2-bench Airline model?

    Not automatically. LongCat-Flash-Thinking-2601 leads tau2-bench Airline, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh tau2-bench Airline against input/output price, context window, and related evals.

    How often is the tau2-bench Airline leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is tau2-bench Airline?

    tau2-bench Airline is a public LLM eval (the tau2-bench Airline score). This page ranks models that have published a score, next to live API prices.

    Where is the tau2-bench Airline leaderboard?

    This page is the tau2-bench Airline leaderboard. Models are sorted by tau2-bench Airline, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this tau2-bench Airline ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published tau2-bench Airline score next to live API $/M so you can pick a production SKU, not only a trophy number.