OSWorld

    🏆 Leaderboard

    As of August 20, 2026, Qwen3.8 Max is #1 for OSWorld at 86.1%. Ranked by OSWorld: can the model drive a real desktop (click, type, finish the task). 38 models in this index have a published OSWorld score. Methodology: OSWorld (https://os-world.github.io/). OSWorld leaderboard with live API prices. OSWorld — benchmarks AI agents on performing real desktop tasks across operating systems. Official methodology: OSWorld (https://os-world.github.io/).

    Updated August 20, 2026282 models33 providers
    Qwen3.8 MaxOSS
    Qwen · Open Source
    86.1%$8.00
    Claude Mythos 5
    Anthropic · Proprietary
    85%$60.00
    Claude Fable 5
    Anthropic · Proprietary
    85%$60.00
    Kimi K3OSS
    Moonshot AI · Open Source
    84.8%$18.00
    Qwen3.8-27BOSS
    Qwen · Open Source
    84.3%$3.65
    Gemini 3.6 Flash
    Google · Proprietary
    83%$4.50
    Claude Sonnet 5
    Anthropic · Proprietary
    81.2%$12.00
    GPT-5.5
    OpenAI · Proprietary
    78.7%$35.00
    Claude Sonnet 4.6
    Anthropic · Proprietary
    78.5%$18.00
    Seed 2.1 Turbo
    ByteDance · Proprietary
    76.4%$3.00
    Gemini 3.5 Flash-Lite
    Google · Proprietary
    74%$2.80
    Claude Opus 4.6
    Anthropic · Proprietary
    72.7%$30.00
    GPT-5.4 mini
    OpenAI · Proprietary
    72.1%$5.25
    Claude Opus 5
    Anthropic · Proprietary
    70.6%$30.00
    Qwen3 VL 235B A22B InstructOSS
    Qwen · Open Source
    66.7%$1.79
    Claude Opus 4.5
    Anthropic · Proprietary
    66.3%$30.00
    GPT-5.4
    OpenAI · Proprietary
    64.9%$17.50
    Kimi K2.5OSS
    Moonshot AI · Open Source
    63.3%$3.68
    GLM-5V-TurboOSS
    Z AI · Open Source
    62.3%$5.20
    Claude Sonnet 4.5
    Anthropic · Proprietary
    61.4%$18.00
    GPT-5.3 Codex
    OpenAI · Proprietary
    59%$15.75
    Claude Haiku 4.5
    Anthropic · Proprietary
    50.7%$6.00
    GPT-5.2 Pro
    OpenAI · Proprietary
    47.8%$189.00
    Gemini 3 Pro
    Google · Proprietary
    47.5%$14.00
    GPT-5.2
    OpenAI · Proprietary
    39.2%$15.75
    GPT-5.4 nano
    OpenAI · Proprietary
    39%$1.45
    Claude Sonnet 4
    Anthropic · Proprietary
    38.6%$18.00
    Qwen3 VL 235B A22B ThinkingOSS
    Qwen · Open Source
    38.1%$3.94
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    35.8%$18.00
    Qwen3 VL 8B ThinkingOSS
    Qwen · Open Source
    33.9%$2.27
    Qwen3 VL 8B InstructOSS
    Qwen · Open Source
    33.9%$0.58
    Qwen3 VL 32B InstructOSS
    Qwen · Open Source · via OpenRouter
    32.6%$0.52
    Qwen3 VL 4B ThinkingOSS
    Qwen · Open Source
    31.4%$1.10
    Qwen3 VL 30B A3B ThinkingOSS
    Qwen · Open Source
    30.6%$1.20
    Qwen3 VL 30B A3B InstructOSS
    Qwen · Open Source
    30.3%$0.90
    Qwen3 VL 4B InstructOSS
    Qwen · Open Source
    26.2%$0.70
    o3
    OpenAI · Proprietary
    23%$10.00
    Qwen2.5 VL 72B InstructOSS
    Qwen · Open Source · via OpenRouter
    8.8%$1.80
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Claude 3 Haiku
    Anthropic · Proprietary
    $1.50
    Claude 3 Opus
    Anthropic · Proprietary
    $90.00
    Claude 3 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Haiku
    Anthropic · Proprietary
    $4.80
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude 3.5 Sonnet
    Anthropic · Proprietary
    $18.00
    Claude Mythos Preview
    Anthropic · Proprietary
    $60.00
    Claude Opus 4
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.1
    Anthropic · Proprietary
    $90.00
    Claude Opus 4.7
    Anthropic · Proprietary
    $30.00
    Claude Opus 4.8
    Anthropic · Proprietary
    $30.00
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads OSWorld right now?

    As of August 20, 2026, Qwen3.8 Max by Qwen is #1 for OSWorld at 86.1%. Ranked by OSWorld: can the model drive a real desktop (click, type, finish the task). This board also tracks OSWorld. Next on the same board: Claude Fable 5 and Claude Mythos 5. This osworld leaderboard ranks models by OSWorld. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OSWorld (https://os-world.github.io/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for osworld. Ranked by OSWorld: can the model drive a real desktop (click, type, finish the task). Input and output are dollars per million tokens.
    RankModelOSWorldInput /MOutput /M
    1Qwen3.8 Max86.1%$2.00$6.00
    2Claude Fable 585%$10.00$50.00
    3Claude Mythos 585%$10.00$50.00
    4Kimi K384.8%$3.00$15.00
    5Qwen3.8-27B84.3%$0.45$3.20
    6Gemini 3.6 Flash83%$0.75$3.75
    7Claude Sonnet 581.2%$2.00$10.00
    8GPT-5.578.7%$5.00$30.00

    OSWorld FAQ

    Who ranks #1 on the OSWorld leaderboard?

    As of August 20, 2026, Qwen3.8 Max by Qwen ranks #1 on OSWorld at 86.1%. API pricing is $2.00/M input and $6.00/M output.

    What are the top models on OSWorld?

    The current OSWorld ranking as of August 20, 2026 is 1. Qwen3.8 Max at 86.1%; 2. Claude Fable 5 at 85%; 3. Claude Mythos 5 at 85%.

    Which osworld model is the cheapest?

    Qwen3 VL 32B Instruct is the cheapest scored model on this osworld leaderboard at $0.10/M input and $0.42/M output ($0.52 blended). Qwen3.8 Max still leads OSWorld at 86.1%.

    Should I always pick the #1 OSWorld model?

    Not automatically. Qwen3.8 Max leads OSWorld, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh OSWorld against input/output price, context window, and related evals.

    How often is the OSWorld leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is OSWorld?

    OSWorld — benchmarks AI agents on performing real desktop tasks across operating systems. This page ranks models that have published a OSWorld score, with live API token prices on the same row. Official methodology: OSWorld (https://os-world.github.io/).

    Where is the OSWorld leaderboard?

    This page is the OSWorld leaderboard. Models are sorted by OSWorld, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this OSWorld ranking different from the official board?

    The official OSWorld page owns the methodology. This page keeps the published OSWorld score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://os-world.github.io/