Long Context

    πŸ† Leaderboard

    As of August 21, 2026, Gemini 3.7 Flash is #1 for long-context retrieval at 97%. Ranked by MRCR: can the model still find planted facts inside a long prompt. Claude vs Gemini context window, ranked by MRCR and Graphwalks. A 1M context window is not the same as usable long context.

    Updated August 21, 2026282 models33 providers
    Gemini 3.7 Flash
    Google Β· Proprietary
    97%β€”β€”β€”$4.50
    Qwen3.8 MaxOSS
    Qwen Β· Open Source
    92.9%β€”β€”β€”$8.00
    Qwen3.7-Plus
    Qwen Β· Proprietary
    91.7%β€”β€”β€”$1.60
    GPT-5.6 Sol
    OpenAI Β· Proprietary
    91.5%β€”β€”90.7%$35.00
    GPT-5.6 Terra
    OpenAI Β· Proprietary
    89.6%β€”β€”76.9%$14.00
    Claude Opus 4.6
    Anthropic Β· Proprietary
    76%β€”β€”61.5%$30.00
    GPT-5.5
    OpenAI Β· Proprietary
    74%β€”β€”45.4%$35.00
    Gemma 4 31BOSS
    Google Β· Open Source
    66.4%β€”β€”β€”$0.54
    Gemini 3.1 Flash-Lite
    Google Β· Proprietary
    60.1%β€”β€”β€”$1.75
    Gemini 3.6 Flash
    Google Β· Proprietary
    54%β€”β€”β€”$4.50
    Gemma 4 26B-A4BOSS
    Google Β· Open Source
    44.1%β€”β€”β€”$0.53
    GPT-5.6 Luna
    OpenAI Β· Proprietary
    41.3%β€”β€”81.3%$1.40
    GPT-5.4 mini
    OpenAI Β· Proprietary
    33.6%β€”β€”76.3%$5.25
    GPT-5.4 nano
    OpenAI Β· Proprietary
    33.1%β€”β€”73.4%$1.45
    Gemini 3.5 Flash
    Google Β· Proprietary
    26.6%β€”β€”β€”$10.50
    Gemini 3.1 Pro
    Google Β· Proprietary
    26.3%β€”β€”β€”$17.50
    Gemini 3 Pro
    Google Β· Proprietary
    26.3%β€”β€”β€”$14.00
    Gemini 3 Flash
    Google Β· Proprietary
    22.1%β€”β€”β€”$3.50
    Gemini 3.5 Flash-Lite
    Google Β· Proprietary
    21.3%β€”β€”β€”$2.80
    Gemini 2.5 Flash-LiteOSS
    Google Β· Open Source
    16.6%β€”β€”β€”$0.50
    Gemini 2.5 Pro Preview 06-05
    Google Β· Proprietary
    16.4%87.5%β€”β€”$11.25
    Gemma 3 27BOSS
    Google Β· Open Source
    13.5%0%β€”β€”$0.30
    ChatGPT-4o Latest
    OpenAI Β· Proprietary
    β€”β€”β€”β€”$12.50
    Claude 3 Haiku
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$1.50
    Claude 3 Opus
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$90.00
    Claude 3 Sonnet
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$18.00
    Claude 3.5 Haiku
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$4.80
    Claude 3.5 Sonnet
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$18.00
    Claude 3.5 Sonnet
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$18.00
    Claude 3.7 Sonnet
    Anthropic Β· Proprietary
    β€”34.4%β€”β€”$18.00
    Claude Fable 5
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$60.00
    Claude Haiku 4.5
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$6.00
    Claude Mythos 5
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$60.00
    Claude Mythos Preview
    Anthropic Β· Proprietary
    β€”β€”β€”80%$60.00
    Claude Opus 4
    Anthropic Β· Proprietary
    β€”37.5%β€”β€”$90.00
    Claude Opus 4.1
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$90.00
    Claude Opus 4.5
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$30.00
    Claude Opus 4.7
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$30.00
    Claude Opus 4.8
    Anthropic Β· Proprietary
    β€”β€”β€”68.1%$30.00
    Claude Opus 5
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$30.00
    Claude Sonnet 4
    Anthropic Β· Proprietary
    β€”36.4%β€”β€”$18.00
    Claude Sonnet 4.5
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$18.00
    Claude Sonnet 4.6
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$18.00
    Claude Sonnet 5
    Anthropic Β· Proprietary
    β€”β€”β€”β€”$12.00
    Command A+OSS
    Cohere Β· Open Source
    β€”β€”β€”β€”$12.50
    Command R+OSS
    Cohere Β· Open Source
    β€”β€”β€”β€”$1.25
    Composer 2
    Cursor Β· Proprietary
    β€”β€”β€”β€”$3.00
    Composer 2 Fast
    Cursor Β· Proprietary
    β€”β€”β€”β€”$9.00
    DeepSeek R1 Distill Llama 70BOSS
    DeepSeek Β· Open Source
    β€”β€”β€”β€”$0.50
    DeepSeek R1 Distill Qwen 32BOSS
    DeepSeek Β· Open Source
    β€”β€”β€”β€”$0.30
    Showing 1–50 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab β€” every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Does a 1 million token context window actually work?

    As of August 21, 2026, Gemini 3.7 Flash by Google is #1 for long-context retrieval at 97%. Ranked by MRCR: can the model still find planted facts inside a long prompt. This board also tracks MRCR v2, Fiction.liveBench, AA-LCR, Graphwalks BFS (0K-128K). Next on the same board: Qwen3.8 Max and Qwen3.7-Plus. Related leaders: o3 on Fiction.liveBench at 100%; Muse Glimmer-30B on AA-LCR at 80%. This long context leaderboard ranks models by MRCR v2. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for long context. Ranked by MRCR: can the model still find planted facts inside a long prompt. Input and output are dollars per million tokens.
    RankModelMRCR v2Input /MOutput /M
    1Gemini 3.7 Flash97%$0.75$3.75
    2Qwen3.8 Max92.9%$2.00$6.00
    3Qwen3.7-Plus91.7%$0.32$1.28
    4GPT-5.6 Sol91.5%$5.00$30.00
    5GPT-5.6 Terra89.6%$2.00$12.00
    6Claude Opus 4.676%$5.00$25.00
    7GPT-5.574%$5.00$30.00
    8Gemma 4 31B66.4%$0.14$0.40

    Claude vs Gemini context window

    People search Claude context window and Gemini context window more than best long context LLM. The listed window is the cap. MRCR and Graphwalks are whether the model can still find needles at 128K and beyond. A 1M window that forgets after 64K is a billing trap.

    What is needle in a haystack for LLMs?

    Needle-in-a-haystack plants a fact in a long prompt and asks for it back. MRCR is the harder multi-needle version. That is the eval to sort if you were burned by a model that β€œsupports 1M” and still misses the clause on page 40.

    Long Context FAQ

    Who ranks #1 on the Long Context leaderboard?

    As of August 21, 2026, Gemini 3.7 Flash by Google ranks #1 on MRCR v2 at 97%. API pricing is $0.75/M input and $3.75/M output.

    Does a 1 million token context window actually work?

    As of August 21, 2026, Gemini 3.7 Flash by Google is #1 for long-context retrieval at 97%. Ranked by MRCR: can the model still find planted facts inside a long prompt. This board also tracks MRCR v2, Fiction.liveBench, AA-LCR, Graphwalks BFS (0K-128K). Next on the same board: Qwen3.8 Max and Qwen3.7-Plus. Related leaders: o3 on Fiction.liveBench at 100%; Muse Glimmer-30B on AA-LCR at 80%.

    What are the top models on MRCR v2?

    The current MRCR v2 ranking as of August 21, 2026 is 1. Gemini 3.7 Flash at 97%; 2. Qwen3.8 Max at 92.9%; 3. Qwen3.7-Plus at 91.7%.

    Which long context model is the cheapest?

    Gemma 3 27B is the cheapest scored model on this long context leaderboard at $0.10/M input and $0.20/M output ($0.30 blended). Gemini 3.7 Flash still leads MRCR v2 at 97%.

    Should I always pick the #1 MRCR v2 model?

    Not automatically. Gemini 3.7 Flash leads MRCR v2, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh MRCR v2 against input/output price, context window, and related evals.

    How often is the Long Context leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 21, 2026. Treat it as a current index, not a one-off blog post.

    What is the best AI for long context?

    A 1M context window is not enough. MRCR and Graphwalks measure whether the model can still retrieve and reason inside that window.

    What is MRCR v2?

    MRCR v2 is OpenAI's multi-needle retrieval eval. It hides several facts in long prompts and checks whether the model can pull them back.

    Should I just pick the largest context window?

    No. Some 1M models lose accuracy after 128K. Sort by MRCR or Graphwalks, then check the listed context window and token price.

    What is Claude's context window?

    Claude's listed context window is on each model row. Listed size is not usable size. MRCR and Graphwalks measure whether the model still retrieves facts at 128K, 256K, and 1M. Sort those columns, then read the sticker context number.

    What is Gemini's context window?

    Gemini advertises very large context, including 1M-class windows. The question people mean is whether retrieval still works at that length. This page ranks MRCR and Graphwalks, not the marketing number.

    What is MRCR?

    MRCR (Multi-Round Coreference / multi-needle retrieval) hides several facts in a long prompt and checks whether the model can pull them back. It is the usual public long-context exam next to needle-in-a-haystack demos.

    Should I use RAG instead of a huge context window?

    Usually yes for production. Filling 1M tokens is slow and expensive. Use long context when the document must stay in-prompt. Use RAG when you can retrieve. Sort MRCR to see if the giant window is real.