Who ranks #1 on the Writing leaderboard?
As of August 21, 2026, Claude Fable 5 by Anthropic ranks #1 on Chatbot Arena at 1506. API pricing is $10.00/M input and $50.00/M output.
As of August 21, 2026, Claude Fable 5 is #1 for writing at 1506. Ranked by Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. Best AI for writing, ranked by Chatbot Arena votes (humans pick the better reply, blind). MMLU-Pro is a knowledge quiz and is not how this page picks a winner.
![]() Claude Fable 5 Anthropic ยท Proprietary | 1506 | 91.5% | โ | 68.3% | 78.3% | โ | โ | โ | โ | โ | โ | $60.00 |
Muse Spark 1.2 Meta ยท Proprietary | 1498 | 88.3% | โ | โ | โ | โ | โ | โ | โ | โ | โ | $5.50 |
![]() Claude Opus 4.6 Anthropic ยท Proprietary | 1497 | โ | โ | 41.0% | 76.3% | โ | 91.1% | โ | โ | โ | โ | $30.00 |
![]() Claude Opus 4.7 Anthropic ยท Proprietary | 1494 | 89.9% | โ | 50.6% | 76.9% | โ | 91.5% | โ | โ | โ | โ | $30.00 |
![]() Claude Opus 5 Anthropic ยท Proprietary | 1493 | 91.6% | โ | 56.7% | โ | โ | โ | โ | โ | โ | โ | $30.00 |
Qwen3.8 MaxOSS Qwen ยท Open Source | 1491 | 88.6% | โ | 46.3% | โ | โ | โ | โ | 82.8% | โ | โ | $8.00 |
Gemini 3.7 Flash Google ยท Proprietary | 1490 | 90.1% | โ | 71.2% | โ | โ | โ | โ | โ | โ | โ | $4.50 |
Muse Spark 1.1 Meta ยท Proprietary | 1489 | 88.7% | โ | โ | โ | โ | โ | โ | โ | โ | โ | $5.50 |
Kimi K3OSS Moonshot AI ยท Open Source | 1489 | 88.0% | โ | 42.7% | โ | โ | โ | โ | โ | โ | โ | $18.00 |
Gemini 3.1 Pro Google ยท Proprietary | 1486 | 91.0% | โ | 77.3% | 79.9% | โ | 92.6% | โ | โ | โ | โ | $17.50 |
Gemini 3 Pro Google ยท Proprietary | 1485 | 90.1% | โ | 72.1% | 73.4% | โ | 91.8% | โ | โ | โ | โ | $14.00 |
Gemini 3.6 Flash Google ยท Proprietary | 1484 | 89.3% | โ | 68.7% | โ | โ | โ | โ | โ | โ | โ | $4.50 |
![]() GPT-5.5 OpenAI ยท Proprietary | 1482 | 88.1% | โ | โ | 80.7% | โ | โ | โ | โ | โ | โ | $35.00 |
![]() GPT-5.6 Sol OpenAI ยท Proprietary | 1481 | 89.1% | โ | 71.6% | โ | โ | โ | โ | โ | โ | โ | $35.00 |
![]() Claude Opus 4.8 Anthropic ยท Proprietary | 1481 | 89.6% | โ | 39.5% | 77.2% | โ | โ | โ | โ | โ | โ | $30.00 |
Gemini 3.5 Flash Google ยท Proprietary | 1477 | 89.5% | โ | 68.4% | 75.0% | โ | โ | โ | โ | โ | โ | $10.50 |
![]() ChatGPT-4o Latest OpenAI ยท Proprietary | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $12.50 |
![]() Claude 3 Haiku Anthropic ยท Proprietary | โ | โ | 75.2% | โ | โ | โ | โ | โ | โ | 85.9% | 74.2% | $1.50 |
![]() Claude 3 Opus Anthropic ยท Proprietary | โ | 68.5% | 86.8% | โ | 49.2% | โ | โ | โ | โ | 95.4% | 88.5% | $90.00 |
![]() Claude 3 Sonnet Anthropic ยท Proprietary | โ | 56.8% | 79% | โ | โ | โ | โ | โ | โ | 89% | 75.1% | $18.00 |
![]() Claude 3.5 Haiku Anthropic ยท Proprietary | โ | 65% | 74.3% | 6.7% | 43.5% | โ | โ | 7.3% | โ | โ | โ | $4.80 |
![]() Claude 3.5 Sonnet Anthropic ยท Proprietary | โ | 77.6% | 90.4% | โ | 59.0% | โ | โ | 8.0% | โ | โ | โ | $18.00 |
![]() Claude 3.5 Sonnet Anthropic ยท Proprietary | โ | 76.1% | 90.4% | โ | โ | โ | โ | โ | โ | โ | โ | $18.00 |
![]() Claude 3.7 Sonnet Anthropic ยท Proprietary | โ | 80.7% | โ | โ | 76.1% | 93.2% | 86.1% | 8.1% | โ | โ | โ | $18.00 |
![]() Claude Haiku 4.5 Anthropic ยท Proprietary | โ | โ | โ | 5.9% | โ | โ | 83% | โ | โ | โ | โ | $6.00 |
![]() Claude Mythos 5 Anthropic ยท Proprietary | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $60.00 |
![]() Claude Mythos Preview Anthropic ยท Proprietary | โ | โ | โ | โ | โ | โ | 92.7% | โ | โ | โ | โ | $60.00 |
![]() Claude Opus 4 Anthropic ยท Proprietary | โ | 86.2% | โ | โ | โ | โ | 88.8% | 8.4% | โ | โ | โ | $90.00 |
![]() Claude Opus 4.1 Anthropic ยท Proprietary | โ | 87.2% | โ | 34.8% | โ | โ | 89.5% | 8.5% | โ | โ | โ | $90.00 |
![]() Claude Opus 4.5 Anthropic ยท Proprietary | โ | 85.6% | โ | 41.8% | 76.0% | โ | 90.8% | โ | โ | โ | โ | $30.00 |
![]() Claude Sonnet 4 Anthropic ยท Proprietary | โ | 79.4% | โ | โ | โ | โ | 86.5% | 8.1% | โ | โ | โ | $18.00 |
![]() Claude Sonnet 4.5 Anthropic ยท Proprietary | โ | โ | โ | 23.6% | โ | โ | 89.1% | โ | โ | โ | โ | $18.00 |
![]() Claude Sonnet 4.6 Anthropic ยท Proprietary | โ | 87.3% | 89.3% | 29% | 75.5% | โ | 89.3% | โ | โ | โ | โ | $18.00 |
![]() Claude Sonnet 5 Anthropic ยท Proprietary | โ | 87.5% | โ | 25% | โ | โ | โ | โ | โ | โ | โ | $12.00 |
Command A+OSS Cohere ยท Open Source | โ | โ | โ | โ | โ | โ | โ | โ | 74% | โ | โ | $12.50 |
Command R+OSS Cohere ยท Open Source | โ | โ | 75.7% | โ | โ | โ | โ | โ | โ | 88.6% | 85.4% | $1.25 |
![]() Composer 2 Cursor ยท Proprietary | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $3.00 |
![]() Composer 2 Fast Cursor ยท Proprietary | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $9.00 |
DeepSeek R1 Distill Llama 70BOSS DeepSeek ยท Open Source | โ | โ | โ | โ | 54.5% | โ | โ | โ | โ | โ | โ | $0.50 |
DeepSeek R1 Distill Qwen 32BOSS DeepSeek ยท Open Source | โ | โ | โ | โ | 45.5% | โ | โ | โ | โ | โ | โ | $0.30 |
DeepSeek-R1OSS DeepSeek ยท Open Source | โ | 83.2% | โ | โ | 71.6% | โ | โ | 8.3% | โ | โ | โ | $2.74 |
DeepSeek-R1-0528OSS DeepSeek ยท Open Source | โ | 85% | โ | 92.3% | โ | โ | โ | 8.2% | โ | โ | โ | $2.74 |
DeepSeek-V2.5OSS DeepSeek ยท Open Source | โ | โ | 80.4% | โ | โ | โ | โ | โ | โ | โ | โ | $0.42 |
DeepSeek-V3OSS DeepSeek ยท Open Source | โ | 75.9% | 88.5% | 24.9% | 60.5% | 86.1% | โ | โ | โ | 88.9% | 85.2% | $1.37 |
DeepSeek-V3 0324OSS DeepSeek ยท Open Source | โ | 81.2% | โ | โ | 66.9% | โ | โ | 7.7% | โ | โ | โ | $1.42 |
DeepSeek-V3.1OSS DeepSeek ยท Open Source | โ | 83.7% | โ | 93.4% | โ | โ | โ | โ | โ | โ | โ | $1.27 |
DeepSeek-V3.2OSS DeepSeek ยท Open Source ยท via OpenRouter | โ | 85% | โ | โ | โ | โ | โ | โ | โ | โ | โ | $0.57 |
DeepSeek-V3.2 (Non-thinking)OSS DeepSeek ยท Open Source | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $0.70 |
DeepSeek-V3.2-ExpOSS DeepSeek ยท Open Source | โ | 85% | โ | 97.1% | โ | โ | โ | โ | โ | โ | โ | $0.68 |
DeepSeek-V4-Flash-0423OSS DeepSeek ยท Open Source | โ | 86.4% | โ | 28.9% | โ | โ | โ | โ | โ | โ | โ | $0.30 |
Next step
Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.
The index
282 models across 33 providers. Search or jump to a lab โ every model page stays linked here.







As of August 21, 2026, Claude Fable 5 by Anthropic is #1 for writing at 1506. Ranked by Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. This board also tracks Chatbot Arena, MMLU-Pro, MMLU, SimpleQA. Next on the same board: Muse Spark 1.2 and Claude Opus 4.6. Related leaders: Claude Opus 5 on MMLU-Pro at 91.6%; GPT-5 on MMLU at 92.5%. This writing leaderboard ranks models by Chatbot Arena. Scores come from public evals. Prices are the live API rates in the table above.
Sources: Chatbot Arena (LMArena) (https://lmarena.ai/); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)
| Rank | Model | Chatbot Arena | Input /M | Output /M |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 1506 | $10.00 | $50.00 |
| 2 | Muse Spark 1.2 | 1498 | $1.25 | $4.25 |
| 3 | Claude Opus 4.6 | 1497 | $5.00 | $25.00 |
| 4 | Claude Opus 4.7 | 1494 | $5.00 | $25.00 |
| 5 | Claude Opus 5 | 1493 | $5.00 | $25.00 |
| 6 | Qwen3.8 Max | 1491 | $2.00 | $6.00 |
| 7 | Gemini 3.7 Flash | 1490 | $0.75 | $3.75 |
| 8 | Kimi K3 | 1489 | $3.00 | $15.00 |
MMLU-Pro is a multiple-choice knowledge exam: law, science, math, and similar. People see it on leaderboards because it is dense, not because it measures writing. A model can ace MMLU-Pro and still write like a press release. This page ranks writing by Chatbot Arena votes: two replies, no names, humans pick a winner.
Search results for best AI for writing are full of app roundups. Those tools wrap a model. This leaderboard ranks the models on preference and language evals, with token price. If you already pay for ChatGPT or Claude, you already have a writer. The question is which SKU to call in production.
Taste tests usually go to Claude. Tight formats and tool-using drafts often go to GPT. MMLU will not tell you that. Sort Arena Elo, read SimpleQA so you do not ship a fluent liar, then check $/M if you generate in bulk.
Rank one eval at a time. All LLM benchmarks.
As of August 21, 2026, Claude Fable 5 by Anthropic ranks #1 on Chatbot Arena at 1506. API pricing is $10.00/M input and $50.00/M output.
As of August 21, 2026, Claude Fable 5 by Anthropic is #1 for writing at 1506. Ranked by Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. This board also tracks Chatbot Arena, MMLU-Pro, MMLU, SimpleQA. Next on the same board: Muse Spark 1.2 and Claude Opus 4.6. Related leaders: Claude Opus 5 on MMLU-Pro at 91.6%; GPT-5 on MMLU at 92.5%.
The current Chatbot Arena ranking as of August 21, 2026 is 1. Claude Fable 5 at 1506; 2. Muse Spark 1.2 at 1498; 3. Claude Opus 4.6 at 1497.
Gemini 3.7 Flash is the cheapest scored model on this writing leaderboard at $0.75/M input and $3.75/M output ($4.50 blended). Claude Fable 5 still leads Chatbot Arena at 1506.
Not automatically. Claude Fable 5 leads Chatbot Arena, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh Chatbot Arena against input/output price, context window, and related evals.
Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 21, 2026. Treat it as a current index, not a one-off blog post.
This page ranks writing by Chatbot Arena votes: humans pick the better reply without seeing the names. That is a taste and prose signal. MMLU-Pro is a multiple-choice knowledge quiz and is not the ranking.
MMLU-Pro is a hard multiple-choice exam (science, law, math, and similar). A high MMLU-Pro score means the model knows facts. It does not mean the model writes well. The writing board still shows the column so you can sort it, but the default rank is Arena votes.
No. MMLU and MMLU-Pro measure knowledge, not prose. HellaSwag is sentence completion. For writing, use Chatbot Arena votes or the Lech Mazur writing eval, then read a sample yourself.
Only if the writing is good enough. Rank by Arena votes first, then check token price. A cheap model that sounds stiff still costs you edits.
Best AI writer usually means a writing app, not a model. This page ranks the models those apps call by Chatbot Arena votes (humans pick the better reply, blind), then shows API price. MMLU-Pro is a knowledge quiz and is not the ranking. Jasper vs ChatGPT is a product comparison.
Creative writing is preference-heavy. Arena Elo and writing-specific evals beat MMLU. Claude often wins taste tests. GPT often wins instruction-following. Sample both. Do not trust a knowledge exam for prose.
Essay help is a mix of structure, citation, and hallucination rate. SimpleQA is the factuality check. MMLU is knowledge, not voice. Rank this board, then read a page of output yourself.
Claude vs ChatGPT writing is the usual split: Claude for tone, GPT for following a tight brief. This table puts preference scores and price on one row so you can stop arguing from anecdotes.
MMLU-Pro is a hard multiple-choice knowledge test. Labs love it because almost every model has a score. It says nothing about whether the prose is any good. This writing board defaults to Chatbot Arena votes. You can still sort the MMLU-Pro column if you want a quiz rank.
No. MMLU is a knowledge test. HellaSwag is sentence completion. Neither is a prose judge. Use Arena Elo and a writing eval, then your own samples.