Writing

    ๐Ÿ† Leaderboard

    As of August 21, 2026, Claude Fable 5 is #1 for writing at 1506. Ranked by Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. Best AI for writing, ranked by Chatbot Arena votes (humans pick the better reply, blind). MMLU-Pro is a knowledge quiz and is not how this page picks a winner.

    Updated August 21, 2026282 models33 providers
    Claude Fable 5
    Anthropic ยท Proprietary
    150691.5%โ€”68.3%78.3%โ€”โ€”โ€”โ€”โ€”โ€”$60.00
    Muse Spark 1.2
    Meta ยท Proprietary
    149888.3%โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$5.50
    Claude Opus 4.6
    Anthropic ยท Proprietary
    1497โ€”โ€”41.0%76.3%โ€”91.1%โ€”โ€”โ€”โ€”$30.00
    Claude Opus 4.7
    Anthropic ยท Proprietary
    149489.9%โ€”50.6%76.9%โ€”91.5%โ€”โ€”โ€”โ€”$30.00
    Claude Opus 5
    Anthropic ยท Proprietary
    149391.6%โ€”56.7%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$30.00
    Qwen3.8 MaxOSS
    Qwen ยท Open Source
    149188.6%โ€”46.3%โ€”โ€”โ€”โ€”82.8%โ€”โ€”$8.00
    Gemini 3.7 Flash
    Google ยท Proprietary
    149090.1%โ€”71.2%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$4.50
    Muse Spark 1.1
    Meta ยท Proprietary
    148988.7%โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$5.50
    Kimi K3OSS
    Moonshot AI ยท Open Source
    148988.0%โ€”42.7%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$18.00
    Gemini 3.1 Pro
    Google ยท Proprietary
    148691.0%โ€”77.3%79.9%โ€”92.6%โ€”โ€”โ€”โ€”$17.50
    Gemini 3 Pro
    Google ยท Proprietary
    148590.1%โ€”72.1%73.4%โ€”91.8%โ€”โ€”โ€”โ€”$14.00
    Gemini 3.6 Flash
    Google ยท Proprietary
    148489.3%โ€”68.7%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$4.50
    GPT-5.5
    OpenAI ยท Proprietary
    148288.1%โ€”โ€”80.7%โ€”โ€”โ€”โ€”โ€”โ€”$35.00
    GPT-5.6 Sol
    OpenAI ยท Proprietary
    148189.1%โ€”71.6%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$35.00
    Claude Opus 4.8
    Anthropic ยท Proprietary
    148189.6%โ€”39.5%77.2%โ€”โ€”โ€”โ€”โ€”โ€”$30.00
    Gemini 3.5 Flash
    Google ยท Proprietary
    147789.5%โ€”68.4%75.0%โ€”โ€”โ€”โ€”โ€”โ€”$10.50
    ChatGPT-4o Latest
    OpenAI ยท Proprietary
    โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$12.50
    Claude 3 Haiku
    Anthropic ยท Proprietary
    โ€”โ€”75.2%โ€”โ€”โ€”โ€”โ€”โ€”85.9%74.2%$1.50
    Claude 3 Opus
    Anthropic ยท Proprietary
    โ€”68.5%86.8%โ€”49.2%โ€”โ€”โ€”โ€”95.4%88.5%$90.00
    Claude 3 Sonnet
    Anthropic ยท Proprietary
    โ€”56.8%79%โ€”โ€”โ€”โ€”โ€”โ€”89%75.1%$18.00
    Claude 3.5 Haiku
    Anthropic ยท Proprietary
    โ€”65%74.3%6.7%43.5%โ€”โ€”7.3%โ€”โ€”โ€”$4.80
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    โ€”77.6%90.4%โ€”59.0%โ€”โ€”8.0%โ€”โ€”โ€”$18.00
    Claude 3.5 Sonnet
    Anthropic ยท Proprietary
    โ€”76.1%90.4%โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$18.00
    Claude 3.7 Sonnet
    Anthropic ยท Proprietary
    โ€”80.7%โ€”โ€”76.1%93.2%86.1%8.1%โ€”โ€”โ€”$18.00
    Claude Haiku 4.5
    Anthropic ยท Proprietary
    โ€”โ€”โ€”5.9%โ€”โ€”83%โ€”โ€”โ€”โ€”$6.00
    Claude Mythos 5
    Anthropic ยท Proprietary
    โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$60.00
    Claude Mythos Preview
    Anthropic ยท Proprietary
    โ€”โ€”โ€”โ€”โ€”โ€”92.7%โ€”โ€”โ€”โ€”$60.00
    Claude Opus 4
    Anthropic ยท Proprietary
    โ€”86.2%โ€”โ€”โ€”โ€”88.8%8.4%โ€”โ€”โ€”$90.00
    Claude Opus 4.1
    Anthropic ยท Proprietary
    โ€”87.2%โ€”34.8%โ€”โ€”89.5%8.5%โ€”โ€”โ€”$90.00
    Claude Opus 4.5
    Anthropic ยท Proprietary
    โ€”85.6%โ€”41.8%76.0%โ€”90.8%โ€”โ€”โ€”โ€”$30.00
    Claude Sonnet 4
    Anthropic ยท Proprietary
    โ€”79.4%โ€”โ€”โ€”โ€”86.5%8.1%โ€”โ€”โ€”$18.00
    Claude Sonnet 4.5
    Anthropic ยท Proprietary
    โ€”โ€”โ€”23.6%โ€”โ€”89.1%โ€”โ€”โ€”โ€”$18.00
    Claude Sonnet 4.6
    Anthropic ยท Proprietary
    โ€”87.3%89.3%29%75.5%โ€”89.3%โ€”โ€”โ€”โ€”$18.00
    Claude Sonnet 5
    Anthropic ยท Proprietary
    โ€”87.5%โ€”25%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$12.00
    Command A+OSS
    Cohere ยท Open Source
    โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”74%โ€”โ€”$12.50
    Command R+OSS
    Cohere ยท Open Source
    โ€”โ€”75.7%โ€”โ€”โ€”โ€”โ€”โ€”88.6%85.4%$1.25
    Composer 2
    Cursor ยท Proprietary
    โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$3.00
    Composer 2 Fast
    Cursor ยท Proprietary
    โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$9.00
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek ยท Open Source
    โ€”โ€”โ€”โ€”54.5%โ€”โ€”โ€”โ€”โ€”โ€”$0.50
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek ยท Open Source
    โ€”โ€”โ€”โ€”45.5%โ€”โ€”โ€”โ€”โ€”โ€”$0.30
    DeepSeek-R1OSS
    DeepSeek ยท Open Source
    โ€”83.2%โ€”โ€”71.6%โ€”โ€”8.3%โ€”โ€”โ€”$2.74
    DeepSeek-R1-0528OSS
    DeepSeek ยท Open Source
    โ€”85%โ€”92.3%โ€”โ€”โ€”8.2%โ€”โ€”โ€”$2.74
    DeepSeek-V2.5OSS
    DeepSeek ยท Open Source
    โ€”โ€”80.4%โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$0.42
    DeepSeek-V3OSS
    DeepSeek ยท Open Source
    โ€”75.9%88.5%24.9%60.5%86.1%โ€”โ€”โ€”88.9%85.2%$1.37
    DeepSeek-V3 0324OSS
    DeepSeek ยท Open Source
    โ€”81.2%โ€”โ€”66.9%โ€”โ€”7.7%โ€”โ€”โ€”$1.42
    DeepSeek-V3.1OSS
    DeepSeek ยท Open Source
    โ€”83.7%โ€”93.4%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$1.27
    DeepSeek-V3.2OSS
    DeepSeek ยท Open Source ยท via OpenRouter
    โ€”85%โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$0.57
    DeepSeek-V3.2 (Non-thinking)OSS
    DeepSeek ยท Open Source
    โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”โ€”$0.70
    DeepSeek-V3.2-ExpOSS
    DeepSeek ยท Open Source
    โ€”85%โ€”97.1%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$0.68
    DeepSeek-V4-Flash-0423OSS
    DeepSeek ยท Open Source
    โ€”86.4%โ€”28.9%โ€”โ€”โ€”โ€”โ€”โ€”โ€”$0.30
    Showing 1โ€“50 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab โ€” every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    What is the best AI for writing?

    As of August 21, 2026, Claude Fable 5 by Anthropic is #1 for writing at 1506. Ranked by Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. This board also tracks Chatbot Arena, MMLU-Pro, MMLU, SimpleQA. Next on the same board: Muse Spark 1.2 and Claude Opus 4.6. Related leaders: Claude Opus 5 on MMLU-Pro at 91.6%; GPT-5 on MMLU at 92.5%. This writing leaderboard ranks models by Chatbot Arena. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: Chatbot Arena (LMArena) (https://lmarena.ai/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for writing. Ranked by Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. Input and output are dollars per million tokens.
    RankModelChatbot ArenaInput /MOutput /M
    1Claude Fable 51506$10.00$50.00
    2Muse Spark 1.21498$1.25$4.25
    3Claude Opus 4.61497$5.00$25.00
    4Claude Opus 4.71494$5.00$25.00
    5Claude Opus 51493$5.00$25.00
    6Qwen3.8 Max1491$2.00$6.00
    7Gemini 3.7 Flash1490$0.75$3.75
    8Kimi K31489$3.00$15.00

    What is MMLU-Pro?

    MMLU-Pro is a multiple-choice knowledge exam: law, science, math, and similar. People see it on leaderboards because it is dense, not because it measures writing. A model can ace MMLU-Pro and still write like a press release. This page ranks writing by Chatbot Arena votes: two replies, no names, humans pick a winner.

    Best AI writer vs best writing model

    Search results for best AI for writing are full of app roundups. Those tools wrap a model. This leaderboard ranks the models on preference and language evals, with token price. If you already pay for ChatGPT or Claude, you already have a writer. The question is which SKU to call in production.

    Claude vs ChatGPT for writing

    Taste tests usually go to Claude. Tight formats and tool-using drafts often go to GPT. MMLU will not tell you that. Sort Arena Elo, read SimpleQA so you do not ship a fluent liar, then check $/M if you generate in bulk.

    Writing FAQ

    Who ranks #1 on the Writing leaderboard?

    As of August 21, 2026, Claude Fable 5 by Anthropic ranks #1 on Chatbot Arena at 1506. API pricing is $10.00/M input and $50.00/M output.

    What is the best AI for writing?

    As of August 21, 2026, Claude Fable 5 by Anthropic is #1 for writing at 1506. Ranked by Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. This board also tracks Chatbot Arena, MMLU-Pro, MMLU, SimpleQA. Next on the same board: Muse Spark 1.2 and Claude Opus 4.6. Related leaders: Claude Opus 5 on MMLU-Pro at 91.6%; GPT-5 on MMLU at 92.5%.

    What are the top models on Chatbot Arena?

    The current Chatbot Arena ranking as of August 21, 2026 is 1. Claude Fable 5 at 1506; 2. Muse Spark 1.2 at 1498; 3. Claude Opus 4.6 at 1497.

    Which writing model is the cheapest?

    Gemini 3.7 Flash is the cheapest scored model on this writing leaderboard at $0.75/M input and $3.75/M output ($4.50 blended). Claude Fable 5 still leads Chatbot Arena at 1506.

    Should I always pick the #1 Chatbot Arena model?

    Not automatically. Claude Fable 5 leads Chatbot Arena, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh Chatbot Arena against input/output price, context window, and related evals.

    How often is the Writing leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 21, 2026. Treat it as a current index, not a one-off blog post.

    What is the best LLM for writing?

    This page ranks writing by Chatbot Arena votes: humans pick the better reply without seeing the names. That is a taste and prose signal. MMLU-Pro is a multiple-choice knowledge quiz and is not the ranking.

    What is MMLU-Pro?

    MMLU-Pro is a hard multiple-choice exam (science, law, math, and similar). A high MMLU-Pro score means the model knows facts. It does not mean the model writes well. The writing board still shows the column so you can sort it, but the default rank is Arena votes.

    Does MMLU measure writing style?

    No. MMLU and MMLU-Pro measure knowledge, not prose. HellaSwag is sentence completion. For writing, use Chatbot Arena votes or the Lech Mazur writing eval, then read a sample yourself.

    Should I pick the cheapest writing model?

    Only if the writing is good enough. Rank by Arena votes first, then check token price. A cheap model that sounds stiff still costs you edits.

    What is the best AI writer?

    Best AI writer usually means a writing app, not a model. This page ranks the models those apps call by Chatbot Arena votes (humans pick the better reply, blind), then shows API price. MMLU-Pro is a knowledge quiz and is not the ranking. Jasper vs ChatGPT is a product comparison.

    What is the best AI for creative writing?

    Creative writing is preference-heavy. Arena Elo and writing-specific evals beat MMLU. Claude often wins taste tests. GPT often wins instruction-following. Sample both. Do not trust a knowledge exam for prose.

    What is the best AI for essays?

    Essay help is a mix of structure, citation, and hallucination rate. SimpleQA is the factuality check. MMLU is knowledge, not voice. Rank this board, then read a page of output yourself.

    Claude vs ChatGPT for writing: which is better?

    Claude vs ChatGPT writing is the usual split: Claude for tone, GPT for following a tight brief. This table puts preference scores and price on one row so you can stop arguing from anecdotes.

    What is MMLU-Pro, and why isn't writing ranked by it?

    MMLU-Pro is a hard multiple-choice knowledge test. Labs love it because almost every model has a score. It says nothing about whether the prose is any good. This writing board defaults to Chatbot Arena votes. You can still sort the MMLU-Pro column if you want a quiz rank.

    Does MMLU measure writing quality?

    No. MMLU is a knowledge test. HellaSwag is sentence completion. Neither is a prose judge. Use Arena Elo and a writing eval, then your own samples.