CharXiv-R

    🏆 Leaderboard

    As of August 20, 2026, Claude Mythos Preview is #1 for CharXiv-R at 93.2%. Ranked by CharXiv: can the model read scientific charts. 47 models in this index have a published CharXiv-R score. CharXiv-R leaderboard with live API prices. CharXiv-R — tests understanding and reasoning about scientific charts and figures from arXiv papers.

    Updated August 20, 2026282 models33 providers
    Claude Mythos Preview
    Anthropic · Proprietary
    93.2%$60.00
    Kimi K3OSS
    Moonshot AI · Open Source
    91.3%$18.00
    Claude Opus 4.7
    Anthropic · Proprietary
    91%$30.00
    Qwen3.8-27BOSS
    Qwen · Open Source
    90.2%$3.65
    Claude Opus 4.8
    Anthropic · Proprietary
    89.9%$30.00
    Gemini 3.6 Flash
    Google · Proprietary
    89.4%$4.50
    Gemini 3.7 Flash
    Google · Proprietary
    88.7%$4.50
    Muse Spark 1.1
    Meta · Proprietary
    88.4%$5.50
    Claude Sonnet 5
    Anthropic · Proprietary
    88.3%$12.00
    Kimi K2.6OSS
    Moonshot AI · Open Source
    86.7%$4.93
    Qwen3.7-Plus
    Qwen · Proprietary
    85.9%$1.60
    Gemini 3.5 Flash
    Google · Proprietary
    84.2%$10.50
    Seed 2.1 Turbo
    ByteDance · Proprietary
    83.6%$3.00
    GPT-5.2
    OpenAI · Proprietary
    82.1%$15.75
    GPT-5.5 Instant
    OpenAI · Proprietary
    81.6%$35.00
    Qwen3.6 Plus
    Qwen · Proprietary
    81.5%$3.50
    Gemini 3 Pro
    Google · Proprietary
    81.4%$14.00
    GPT-5
    OpenAI · Proprietary
    81.1%$11.25
    MiMo-V2.5OSS
    Xiaomi · Open Source
    81%$0.50
    Gemini 3 Flash
    Google · Proprietary
    80.3%$3.50
    Qwen3.5-27BOSS
    Qwen · Open Source
    79.5%$2.70
    Muse Glimmer-30BOSS
    Meta · Open Source
    78.8%$1.85
    o3
    OpenAI · Proprietary
    78.6%$10.00
    Qwen3.6-27BOSS
    Qwen · Open Source
    78.4%$4.20
    Qwen3.6-35B-A3BOSS
    Qwen · Open Source · via OpenRouter
    78%$1.14
    Qwen3.5-35B-A3BOSS
    Qwen · Open Source
    77.5%$2.25
    Kimi K2.5OSS
    Moonshot AI · Open Source
    77.5%$3.68
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    77.4%$1.50
    Claude Opus 4.6
    Anthropic · Proprietary
    77.4%$30.00
    Qwen3.5-122B-A10BOSS
    Qwen · Open Source
    77.2%$3.60
    Gemini 3.5 Flash-Lite
    Google · Proprietary
    76.5%$2.80
    Gemini 3.1 Flash-Lite
    Google · Proprietary
    73.2%$1.75
    o4-mini
    OpenAI · Proprietary
    72%$5.50
    Qwen3 VL 235B A22B ThinkingOSS
    Qwen · Open Source
    66.1%$3.94
    Qwen3 VL 32B InstructOSS
    Qwen · Open Source · via OpenRouter
    62.8%$0.52
    Qwen3 VL 235B A22B InstructOSS
    Qwen · Open Source
    62.1%$1.79
    GPT-4o
    OpenAI · Proprietary
    58.8%$12.50
    GPT-4.1 mini
    OpenAI · Proprietary
    56.8%$2.00
    GPT-4.1
    OpenAI · Proprietary
    56.7%$10.00
    Qwen3 VL 30B A3B ThinkingOSS
    Qwen · Open Source
    56.6%$1.20
    Qwen3 VL 8B ThinkingOSS
    Qwen · Open Source
    53%$2.27
    Command A+OSS
    Cohere · Open Source
    52.7%$12.50
    Qwen3 VL 4B ThinkingOSS
    Qwen · Open Source
    50.3%$1.10
    Qwen3 VL 30B A3B InstructOSS
    Qwen · Open Source
    48.9%$0.90
    Qwen3 VL 8B InstructOSS
    Qwen · Open Source
    46.4%$0.58
    GPT-4.1 nano
    OpenAI · Proprietary
    40.5%$0.50
    Qwen3 VL 4B InstructOSS
    Qwen · Open Source
    39.7%$0.70
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Claude 3 Haiku
    Anthropic · Proprietary
    $1.50
    Claude 3 Opus
    Anthropic · Proprietary
    $90.00
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads CharXiv-R right now?

    As of August 20, 2026, Claude Mythos Preview by Anthropic is #1 for CharXiv-R at 93.2%. Ranked by CharXiv: can the model read scientific charts. This board also tracks CharXiv-R. Next on the same board: Kimi K3 and Claude Opus 4.7. This charxiv-r leaderboard ranks models by CharXiv-R. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for charxiv-r. Ranked by CharXiv: can the model read scientific charts. Input and output are dollars per million tokens.
    RankModelCharXiv-RInput /MOutput /M
    1Claude Mythos Preview93.2%$10.00$50.00
    2Kimi K391.3%$3.00$15.00
    3Claude Opus 4.791%$5.00$25.00
    4Qwen3.8-27B90.2%$0.45$3.20
    5Claude Opus 4.889.9%$5.00$25.00
    6Gemini 3.6 Flash89.4%$0.75$3.75
    7Gemini 3.7 Flash88.7%$0.75$3.75
    8Muse Spark 1.188.4%$1.25$4.25

    CharXiv-R FAQ

    Who ranks #1 on the CharXiv-R leaderboard?

    As of August 20, 2026, Claude Mythos Preview by Anthropic ranks #1 on CharXiv-R at 93.2%. API pricing is $10.00/M input and $50.00/M output.

    What are the top models on CharXiv-R?

    The current CharXiv-R ranking as of August 20, 2026 is 1. Claude Mythos Preview at 93.2%; 2. Kimi K3 at 91.3%; 3. Claude Opus 4.7 at 91%.

    Which charxiv-r model is the cheapest?

    GPT-4.1 nano is the cheapest scored model on this charxiv-r leaderboard at $0.10/M input and $0.40/M output ($0.50 blended). Claude Mythos Preview still leads CharXiv-R at 93.2%.

    Should I always pick the #1 CharXiv-R model?

    Not automatically. Claude Mythos Preview leads CharXiv-R, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh CharXiv-R against input/output price, context window, and related evals.

    How often is the CharXiv-R leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is CharXiv-R?

    CharXiv-R — tests understanding and reasoning about scientific charts and figures from arXiv papers. This page ranks models that have published a CharXiv-R score, with live API token prices on the same row.

    Where is the CharXiv-R leaderboard?

    This page is the CharXiv-R leaderboard. Models are sorted by CharXiv-R, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this CharXiv-R ranking different from the official board?

    Official eval pages own the methodology. This page keeps the published CharXiv-R score next to live API $/M so you can pick a production SKU, not only a trophy number.