IFEval

    🏆 Leaderboard

    As of August 20, 2026, Qwen3.5-27B is #1 for IFEval at 95%. Ranked by IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.' 49 models in this index have a published IFEval score. Methodology: IFEval (https://github.com/google-research/google-research/tree/master/instruction_following_eval). IFEval leaderboard: rank models by IFEval next to live API token prices. Official methodology: IFEval (https://github.com/google-research/google-research/tree/master/instruction_following_eval).

    Updated August 20, 2026282 models33 providers
    Qwen3.5-27BOSS
    Qwen · Open Source
    95%$2.70
    Qwen3.7-Plus
    Qwen · Proprietary
    94.6%$1.60
    Qwen3.7 Max
    Qwen · Proprietary
    94.3%$5.00
    Qwen3.6 Plus
    Qwen · Proprietary
    94.3%$3.50
    o3-mini
    OpenAI · Proprietary
    93.9%$5.50
    Qwen3.5-122B-A10BOSS
    Qwen · Open Source
    93.4%$3.60
    Claude 3.7 Sonnet
    Anthropic · Proprietary
    93.2%$18.00
    Qwen3.5-397B-A17BOSS
    Qwen · Open Source
    92.6%$4.20
    Nova Pro
    Amazon · Proprietary
    92.1%$4.00
    Llama 3.3 70B InstructOSS
    Meta · Open Source
    92.1%$0.40
    Qwen3.5-35B-A3BOSS
    Qwen · Open Source
    91.9%$2.25
    Qwen3.5-9BOSS
    Qwen · Open Source · via OpenRouter
    91.5%$0.25
    Gemma 3 27BOSS
    Google · Open Source
    90.4%$0.30
    Nemotron Nano 9B v2OSS
    NVIDIA · Open Source · via OpenRouter
    90.3%
    Gemma 3 4BOSS
    Google · Open Source
    90.2%$0.06
    Kimi K2-Instruct-0905OSS
    Moonshot AI · Open Source · via OpenRouter
    89.8%$3.10
    Kimi K2 InstructOSS
    Moonshot AI · Open Source
    89.8%$1.00
    Nova Lite
    Amazon · Proprietary
    89.7%$0.30
    LongCat-Flash-ChatOSS
    Meituan · Open Source
    89.7%$1.50
    Qwen3-Next-80B-A3B-ThinkingOSS
    Qwen · Open Source
    88.9%$1.65
    Gemma 3 12BOSS
    Google · Open Source
    88.9%$0.15
    Qwen3-235B-A22B-Instruct-2507OSS
    Qwen · Open Source
    88.7%$0.95
    Llama 3.1 405B InstructOSS
    Meta · Open Source
    88.6%$1.78
    Qwen3 VL 235B A22B ThinkingOSS
    Qwen · Open Source
    88.2%$3.94
    Qwen3-235B-A22B-Thinking-2507OSS
    Qwen · Open Source
    87.8%$3.30
    Qwen3 VL 235B A22B InstructOSS
    Qwen · Open Source
    87.8%$1.79
    Qwen3-Next-80B-A3B-InstructOSS
    Qwen · Open Source
    87.6%$1.65
    Llama 3.1 70B InstructOSS
    Meta · Open Source
    87.5%$0.40
    GPT-4.1
    OpenAI · Proprietary
    87.4%$10.00
    Nova Micro
    Amazon · Proprietary
    87.2%$0.17
    DeepSeek-V3OSS
    DeepSeek · Open Source
    86.1%$1.37
    Qwen3 VL 30B A3B InstructOSS
    Qwen · Open Source
    85.8%$0.90
    Qwen3 VL 32B InstructOSS
    Qwen · Open Source · via OpenRouter
    84.7%$0.52
    Qwen2.5 72B InstructOSS
    Qwen · Open Source
    84.1%$0.75
    GPT-4.1 mini
    OpenAI · Proprietary
    84.1%$2.00
    Qwen3 VL 8B InstructOSS
    Qwen · Open Source
    83.7%$0.58
    Qwen3 VL 8B ThinkingOSS
    Qwen · Open Source
    83.2%$2.27
    Mistral Small 3 24B InstructOSS
    Mistral · Open Source · via Mistral AI
    82.9%$0.21
    Qwen3 VL 4B ThinkingOSS
    Qwen · Open Source
    82.6%$1.10
    Qwen3 VL 4B InstructOSS
    Qwen · Open Source
    82.3%$0.70
    Qwen3 VL 30B A3B ThinkingOSS
    Qwen · Open Source
    81.7%$1.20
    GPT-4o
    OpenAI · Proprietary
    81%$12.50
    Llama 3.1 8B InstructOSS
    Meta · Open Source
    80.4%$0.06
    Llama 3.2 3B InstructOSS
    Meta · Open Source
    77.4%$0.03lowest
    Granite 3.3 8B InstructOSS
    IBM · Open Source
    74.8%$1.00
    GPT-4.1 nano
    OpenAI · Proprietary
    74.5%$0.50
    Qwen2.5 7B InstructOSS
    Qwen · Open Source
    71.2%$0.60
    Phi 4OSS
    Microsoft · Open Source
    63%$0.21
    Pixtral-12BOSS
    Mistral · Open Source · via Mistral AI
    61.3%$0.30
    ChatGPT-4o Latest
    OpenAI · Proprietary
    $12.50
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All Large Language Models

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Baidu

    1 models

    Inception

    1 models

    inclusionAI

    1 models

    LG AI Research

    1 models

    Nous Research

    1 models

    Sakana AI

    1 models

    StepFun

    1 models

    Tencent

    1 models

    Thinking Machines

    1 models

    Unisound

    1 models

    Upstage

    1 models

    Which model leads IFEval right now?

    As of August 20, 2026, Qwen3.5-27B by Qwen is #1 for IFEval at 95%. Ranked by IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.' This board also tracks IFEval. Next on the same board: Qwen3.7-Plus and Qwen3.7 Max. This ifeval leaderboard ranks models by IFEval. Scores come from public evals. Prices are the live API rates in the table above.

    Sources: IFEval (https://github.com/google-research/google-research/tree/master/instruction_following_eval); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    Top 8 for ifeval. Ranked by IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.' Input and output are dollars per million tokens.
    RankModelIFEvalInput /MOutput /M
    1Qwen3.5-27B95%$0.30$2.40
    2Qwen3.7-Plus94.6%$0.32$1.28
    3Qwen3.7 Max94.3%$1.25$3.75
    4Qwen3.6 Plus94.3%$0.50$3.00
    5o3-mini93.9%$1.10$4.40
    6Qwen3.5-122B-A10B93.4%$0.40$3.20
    7Claude 3.7 Sonnet93.2%$3.00$15.00
    8Qwen3.5-397B-A17B92.6%$0.60$3.60

    IFEval FAQ

    Who ranks #1 on the IFEval leaderboard?

    As of August 20, 2026, Qwen3.5-27B by Qwen ranks #1 on IFEval at 95%. API pricing is $0.30/M input and $2.40/M output.

    What are the top models on IFEval?

    The current IFEval ranking as of August 20, 2026 is 1. Qwen3.5-27B at 95%; 2. Qwen3.7-Plus at 94.6%; 3. Qwen3.7 Max at 94.3%.

    Which ifeval model is the cheapest?

    Llama 3.2 3B Instruct is the cheapest scored model on this ifeval leaderboard at $0.01/M input and $0.02/M output ($0.03 blended). Qwen3.5-27B still leads IFEval at 95%.

    Should I always pick the #1 IFEval model?

    Not automatically. Qwen3.5-27B leads IFEval, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh IFEval against input/output price, context window, and related evals.

    How often is the IFEval leaderboard updated?

    Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 20, 2026. Treat it as a current index, not a one-off blog post.

    What is IFEval?

    IFEval is a public LLM eval (IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.'). This page ranks models that have published a score, next to live API prices. Official methodology: IFEval (https://github.com/google-research/google-research/tree/master/instruction_following_eval).

    Where is the IFEval leaderboard?

    This page is the IFEval leaderboard. Models are sorted by IFEval, with input and output token prices on the same row so you can weigh score against cost. Official boards often omit price; that comparison is the point of this index.

    How is this IFEval ranking different from the official board?

    The official IFEval page owns the methodology. This page keeps the published IFEval score next to live API $/M so you can pick a production SKU, not only a trophy number. Source: https://github.com/google-research/google-research/tree/master/instruction_following_eval