Agentic

    🏆 Leaderboard

    As of August 22, 2026, Kimi K3 is #1 for agent and computer-use work at 91.2%. Ranked by the BrowseComp score Best AI agent ranked by OSWorld, Claude computer use, tau-bench, and tool-use evals. Compare computer-use scores next to API price.

    Updated August 22, 2026282 models33 providers
    Kimi K3OSS
    Moonshot AI · Open Source
    91.2%37.6%84.8%76.5%84.2%$18.00
    Claude Opus 5
    Anthropic · Proprietary
    90.8%43.5%70.6%80.6%85.8%83.4%60.4%70.6%$30.00
    GPT-5.6 Sol
    OpenAI · Proprietary
    90.4%39.9%58%54.1%62.6%52.7%$35.00
    GPT-5.5 Pro
    OpenAI · Proprietary
    90.1%$540.00
    Claude Mythos 5
    Anthropic · Proprietary
    88%85%61.7%85%$60.00
    GPT-5.6 Terra
    OpenAI · Proprietary
    87.5%53.1%60.6%50.2%50.4%$14.00
    Claude Mythos Preview
    Anthropic · Proprietary
    86.9%79.6%$60.00
    Kimi K2.6OSS
    Moonshot AI · Open Source
    86.3%27.9%50%73.1%4.6%$4.93
    Gemini 3.1 Pro
    Google · Proprietary
    85.9%33.5%69.2%78.3%$17.50
    Seed 2.1 Turbo
    ByteDance · Proprietary
    84.9%29.2%76.4%49.1%80.3%$3.00
    Claude Sonnet 5
    Anthropic · Proprietary
    84.7%32.5%81.2%74.7%81.2%46.5%$12.00
    GPT-5.5
    OpenAI · Proprietary
    84.4%38.5%78.7%55.6%75.3%78.7%98%62.2%54.0%13%$35.00
    Claude Opus 4.8
    Anthropic · Proprietary
    84.3%42.5%59.9%82.2%83.4%59.2%87.9%50.2%20.6%$30.00
    Hy3OSS
    Tencent · Open Source
    84.2%25.6%48.5%79.1%55.3%$0.66
    Claude Opus 4.6
    Anthropic · Proprietary
    84%32.4%91.9%72.7%62.7%99.3%55.3%$30.00
    MiniMax M3OSS
    MiniMax · Open Source
    83.5%27.7%74.2%70.1%51.5%4.6%$1.50
    DeepSeek-V4-Pro-MaxOSS
    DeepSeek · Open Source
    83.4%51.8%73.6%$5.22
    GPT-5.6 Luna
    OpenAI · Proprietary
    83.3%53.4%60.5%45.6%50.3%$1.40
    GPT-5.4
    OpenAI · Proprietary
    82.7%36%64.9%54.6%67.2%75%71%51.7%35.1%$17.50
    GLM-5.1OSS
    Z AI · Open Source
    79.3%40.7%71.8%$5.80
    Claude Opus 4.7
    Anthropic · Proprietary
    79.3%33.9%77.3%78%18.2%$30.00
    GPT-5.2 Pro
    OpenAI · Proprietary
    77.9%47.8%77.8%$189.00
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    77.4%54.4%79.6%33.6%$1.50
    Seed 2.0 Pro
    ByteDance · Proprietary
    77.3%$3.50
    MiniMax M2.5OSS
    MiniMax · Open Source
    76.3%6.2%$1.50
    GLM-5OSS
    Z AI · Open Source
    75.9%17.2%67.8%$4.20
    Kimi K2.5OSS
    Moonshot AI · Open Source
    74.9%14.4%63.3%63.3%$3.68
    Claude Sonnet 4.6
    Anthropic · Proprietary
    74.7%23.7%91.7%78.5%61.3%72.1%97.9%49.0%54.9%9.3%$18.00
    DeepSeek-V4-Flash-MaxOSS
    DeepSeek · Open Source
    73.2%47.8%69%$0.42
    Step-3.5-FlashOSS
    StepFun · Open Source
    69%$0.50
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    69%38.3%$4.20
    GPT-5.2
    OpenAI · Proprietary
    65.8%34.4%82%55.9%39.2%46.3%60.6%98.7%86.3%41.1%$15.75
    Qwen3.5-122B-A10BOSS
    Qwen · Open Source
    63.8%58%70.4%$3.60
    MiniMax M2.1OSS
    MiniMax · Open Source
    62%43.5%87%$1.50
    Qwen3.5-35B-A3BOSS
    Qwen · Open Source
    61%54.5%68.6%$2.25
    Qwen3.5-27BOSS
    Qwen · Open Source
    61%56.2%70.3%$2.70
    Kimi K2-Thinking-0905OSS
    Moonshot AI · Open Source
    60.2%$2.47
    MiMo-V2-FlashOSS
    Xiaomi · Open Source
    58.3%$0.40
    LongCat-Flash-Thinking-2601OSS
    Meituan · Open Source
    56.6%88.6%99.3%$1.50
    GPT-5
    OpenAI · Proprietary
    54.9%18.3%81.1%53.6%96.7%55.1%$11.25
    DeepSeek-V4-Flash-0423OSS
    DeepSeek · Open Source
    53.5%43.5%67.4%$0.30
    GLM-4.7OSS
    Z AI · Open Source
    52%8.7%$2.80
    o4-mini
    OpenAI · Proprietary
    51.5%71.8%53.2%$5.50
    DeepSeek-V3.2OSS
    DeepSeek · Open Source · via OpenRouter
    51.4%7%35.2%$0.57
    o3
    OpenAI · Proprietary
    49.7%17.2%80.2%63.0%23%23%58.2%46.6%$10.00
    Solar Pro 4
    Upstage · Proprietary
    49.2%$1.50
    Mistral Medium 3.5OSS
    Mistral · Open Source · via Mistral AI
    48.6%$9.00
    GLM-4.6OSS
    Z AI · Open Source
    45.1%4%72.4%$2.80
    Grok 4 Fast
    xAI · Proprietary
    44.9%$0.70
    Nemotron 3 Ultra (550B A55B)OSS
    NVIDIA · Open Source · via OpenRouter
    44.4%$2.70
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    What is the best AI agent?

    As of August 22, 2026, Kimi K3 by Moonshot AI is #1 for agent and computer-use work at 91.2%. Ranked by the BrowseComp score This board also tracks BrowseComp, APEX-Agents, tau-bench Retail, BFCL. Next on the same board: Claude Opus 5 and GPT-5.6 Sol. Related leaders: Grok 4.6 on APEX-Agents at 57.5%; Claude Opus 4.6 on tau-bench Retail at 91.9%. This agentic leaderboard ranks models by BrowseComp. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: BrowseComp (OpenAI) (https://openai.com/index/browsecomp/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for agentic. Ranked by the BrowseComp score Input and output are dollars per million tokens.
    RankModelBrowseCompInput /MOutput /M
    1Kimi K391.2%$3.00$15.00
    2Claude Opus 590.8%$5.00$25.00
    3GPT-5.6 Sol90.4%$5.00$30.00
    4GPT-5.5 Pro90.1%$60.00$480.00
    5Claude Mythos 588%$10.00$50.00
    6GPT-5.6 Terra87.5%$2.00$12.00
    7Claude Mythos Preview86.9%$10.00$50.00
    8Kimi K2.686.3%$0.96$3.97

    What is Claude computer use?

    Claude computer use is the Anthropic name for letting Claude drive a computer: screenshot in, mouse and keyboard out. People search it more than “best LLM for agents.” OSWorld is the public exam for that loop. A high chat score does not mean the model can finish a desktop task.

    OSWorld vs tau-bench

    OSWorld is GUI computer use. tau-bench is multi-turn tool calls in a text workflow. BrowseComp is web research. Sort the column that matches the loop you ship. Mixing them into one “best agent” trophy is how teams buy the wrong SKU.

    Agentic FAQ

    Who ranks #1 on the Agentic leaderboard?

    As of August 22, 2026, Kimi K3 by Moonshot AI ranks #1 on BrowseComp at 91.2%. API pricing is $3.00/M input and $15.00/M output.

    What is the best AI agent?

    As of August 22, 2026, Kimi K3 by Moonshot AI is #1 for agent and computer-use work at 91.2%. Ranked by the BrowseComp score This board also tracks BrowseComp, APEX-Agents, tau-bench Retail, BFCL. Next on the same board: Claude Opus 5 and GPT-5.6 Sol. Related leaders: Grok 4.6 on APEX-Agents at 57.5%; Claude Opus 4.6 on tau-bench Retail at 91.9%.

    What are the top models on BrowseComp?

    The current BrowseComp ranking as of August 22, 2026 is 1. Kimi K3 at 91.2%; 2. Claude Opus 5 at 90.8%; 3. GPT-5.6 Sol at 90.4%.

    Which agentic model is the cheapest?

    Nemotron 3.5 Lightning (30B A3B) is the cheapest scored model on this agentic leaderboard at $0.05/M input and $0.20/M output ($0.25 blended). Kimi K3 still leads BrowseComp at 91.2%.

    Should I always pick the #1 BrowseComp model?

    Not automatically. Kimi K3 leads BrowseComp, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh BrowseComp against input/output price, context window, and related evals.

    How often is the Agentic leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 22, 2026. Treat it as a current index, not a one-off blog post.

    What is the best AI agent model?

    OSWorld ranks desktop computer-use, tau-bench ranks multi-turn tool conversations, and BrowseComp ranks web research. Sort by the agent loop you ship.

    What is OSWorld?

    OSWorld is a computer-use benchmark. Models have to click, type, and finish real desktop tasks instead of answering a chat prompt.

    What is tau-bench?

    tau-bench and tau2-bench score agents on retail, airline, and telecom workflows that need reliable tool calls over many turns.

    What is Claude computer use?

    Claude computer use is Anthropic's desktop-control loop: the model sees the screen, then clicks and types. OSWorld is the usual public score for that job. Computer use is not the same as a chat model with function calling.

    What is computer use in AI?

    Computer use means the agent operates a real desktop or browser: mouse, keyboard, apps. OSWorld and OSWorld-Verified score that. Tool-calling benches like tau-bench score API tools in a conversation, not a GUI.

    Is the best AI agent the same as the best chatbot?

    No. Chat quality and agent success diverge. A model can win Arena Elo and fail OSWorld. Rank this board for agents, the overall board for chat, and coding for SWE-bench.