LLM Pricing

    Newest first

    Newest models first. Compare 2026 OpenAI, Claude, Gemini, and 200+ LLM token costs, sorted by release date.

    Updated August 22, 2026282 models33 providers
    Qwen3.8-27BOSS
    Qwen · Open Source
    $0.45$3.20262.1K2026-08-14$3.65
    GLM-5.3OSS
    Z AI · Open Source
    $1.40$4.401M2026-08-14$5.80
    Gemini 3.7 Flash
    Google · Proprietary
    $0.75$3.751M2026-08-13$4.50
    DeepSeek-V4-Pro-0813OSS
    DeepSeek · Open Source
    $1.32$3.961M2026-08-13$5.28
    Qwen3.8-2.4T-A95BOSS
    Qwen · Open Source
    $2.00$6.001M2026-08-12$8.00
    Grok 4.6
    xAI · Proprietary
    $2.00$6.00500K2026-08-12$8.00
    Nemotron 3.5 Lightning (30B A3B)OSS
    NVIDIA · Open Source
    $0.05$0.20262.1K2026-08-11$0.25
    MAI-Code-1.1-Flash
    Microsoft · Proprietary
    $0.20$1.20256K2026-08-11$1.40
    Muse Glimmer-30BOSS
    Meta · Open Source
    $0.35$1.50131.1K2026-08-10$1.85
    GPT-5.6 Cyber
    OpenAI · Proprietary
    $12.50$75.00400K2026-08-10$87.50
    Solar Pro 4
    Upstage · Proprietary
    $0.30$1.20524.3K2026-08-06$1.50
    Muse Spark 1.2
    Meta · Proprietary
    $1.25$4.251M2026-08-05$5.50
    Sakana Namazu
    Sakana AI · Proprietary
    $0.95$4.00256K2026-08-03$4.95
    Qwen3.8 MaxOSS
    Qwen · Open Source
    $2.00$6.001M2026-08-02$8.00
    DeepSeek-V4-Flash-0731OSS
    DeepSeek · Open Source
    $0.44$1.321M2026-07-31$1.76
    Inkling-SmallOSS
    Thinking Machines · Open Source · via Thinking Machines Lab
    $0.30$1.20256K2026-07-30$1.50
    Qwen3.7 Flash
    Qwen · Proprietary
    $0.03$0.131M2026-07-27$0.16
    Claude Opus 5
    Anthropic · Proprietary
    $5.00$25.001M2026-07-24$30.00
    Ling-3.0-flash
    inclusionAI · Proprietary
    $0.06$0.18262.1K2026-07-23$0.24
    Laguna S 2.1OSS
    Poolside · Open Source
    $0.10$0.201M2026-07-21$0.30
    Gemini 3.6 Flash
    Google · Proprietary
    $0.75$3.751M2026-07-21$4.50
    Gemini 3.5 Flash-Lite
    Google · Proprietary
    $0.30$2.501M2026-07-21$2.80
    Kimi K3OSS
    Moonshot AI · Open Source
    $3.00$15.001M2026-07-16$18.00
    Grok 4.5
    xAI · Proprietary
    $2.00$6.00500K2026-07-16$8.00
    Muse Spark 1.1
    Meta · Proprietary
    $1.25$4.251M2026-07-09$5.50
    GPT-5.6 Terra
    OpenAI · Proprietary
    $2.00$12.001.1M2026-07-09$14.00
    GPT-5.6 Sol
    OpenAI · Proprietary
    $5.00$30.001.1M2026-07-09$35.00
    GPT-5.6 Luna
    OpenAI · Proprietary
    $0.20$1.201.1M2026-07-09$1.40
    Hy3OSS
    Tencent · Open Source
    $0.13$0.53262.1K2026-07-06$0.66
    Laguna XS 2.1OSS
    Poolside · Open Source
    $0.10$0.20262.1K2026-07-02$0.30
    Claude Sonnet 5
    Anthropic · Proprietary
    $2.00$10.001M2026-06-30$12.00
    Seed 2.1 Turbo
    ByteDance · Proprietary
    $0.50$2.50262.1K2026-06-24$3.00
    GLM-5.2OSS
    Z AI · Open Source
    $1.40$4.401M2026-06-16$5.80
    Kimi K2.7 Code HighSpeedOSS
    Moonshot AI · Open Source
    $1.91$7.94262.1K2026-06-12$9.85
    Kimi K2.7 CodeOSS
    Moonshot AI · Open Source
    $0.96$3.97262.1K2026-06-12$4.93
    North Mini Code 1.0OSS
    Cohere · Open Source · via OpenRouter
    256K2026-06-09
    Claude Mythos 5
    Anthropic · Proprietary
    $10.00$50.001M2026-06-09$60.00
    Claude Fable 5
    Anthropic · Proprietary
    $10.00$50.001M2026-06-09$60.00
    U2
    Unisound · Proprietary
    $0.15$0.302026-06-05$0.45
    Nemotron 3 Ultra (550B A55B)OSS
    NVIDIA · Open Source · via OpenRouter
    $0.50$2.201M2026-06-04$2.70
    MiniMax M3OSS
    MiniMax · Open Source
    $0.30$1.201M2026-06-01$1.50
    Qwen3.7-Plus
    Qwen · Proprietary
    $0.32$1.281M2026-05-31$1.60
    Claude Opus 4.8
    Anthropic · Proprietary
    $5.00$25.001M2026-05-28$30.00
    Command A+OSS
    Cohere · Open Source
    $2.50$10.002026-05-20$12.50
    Qwen3.7 Max
    Qwen · Proprietary
    $1.25$3.751M2026-05-19$5.00
    Gemini 3.5 Flash
    Google · Proprietary
    $1.50$9.001M2026-05-19$10.50
    Grok 4.3
    xAI · Proprietary
    $1.25$2.501M2026-05-06$3.75
    GPT-5.5 Instant
    OpenAI · Proprietary
    $5.00$30.00400K2026-05-05$35.00
    Mistral Medium 3.5OSS
    Mistral · Open Source · via Mistral AI
    $1.50$7.50256K2026-04-29$9.00
    MiMo-V2.5-ProOSS
    Xiaomi · Open Source
    $0.43$0.871M2026-04-27$1.30
    Showing 150 of 282 models

    Next step

    You found the model. Now ship the product.

    Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.

    The index

    All LLM API pricing pages

    282 models across 33 providers. Search or jump to a lab — every model page stays linked here.

    Cursor API pricing

    2 models

    AI21 Labs API pricing

    2 models

    Baidu API pricing

    1 models

    ByteDance API pricing

    2 models

    Cohere API pricing

    3 models

    IBM API pricing

    1 models

    Inception API pricing

    1 models

    inclusionAI API pricing

    1 models

    LG AI Research API pricing

    1 models

    Nous Research API pricing

    1 models

    Poolside API pricing

    2 models

    Sakana AI API pricing

    1 models

    StepFun API pricing

    1 models

    Tencent API pricing

    1 models

    Thinking Machines API pricing

    1 models

    Unisound API pricing

    1 models

    Upstage API pricing

    1 models

    Xiaomi API pricing

    3 models

    LLM API pricing, compared

    As of August 22, 2026, Gemma 3 4B by Google is the cheapest LLM in this index at $0.02/M input and $0.04/M output. This LLM pricing comparison tracks live API costs for 282+ models from OpenAI, Anthropic, Google, DeepSeek, xAI, and other labs, with the LLM leaderboard on the same page.

    OpenAI API pricing, Claude API pricing, and Gemini rates change often. The table above is updated daily from published rates, then normalized to dollars per million tokens so GPT, Claude, Gemini, and open-source models sit on the same scale. AnotherWrapper does not set these rates; it indexes the public list prices.

    Sources: OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)

    LLM API cost calculator

    Default workload: 100k in / 20k out, times 1,000 requests. Cost is list rate × tokens ÷ 1,000,000. Cached input and batch APIs can come in lower than these public rates.

    Typical request: 100K in / 20K out, × 1,000 requests. Input and output are dollars per million tokens.
    ModelInput /MOutput /MPer request1,000 requests
    Gemma 3 4BGoogle$0.02$0.04$0.003$2.80
    GPT-5 nanoOpenAI$0.05$0.40$0.01$13.00
    Claude Haiku 4.5Anthropic$1.00$5.00$0.20$200.00
    Gemini 2.5 Flash-LiteGoogle$0.10$0.40$0.02$18.00
    DeepSeek-V4-Flash-0423DeepSeek$0.10$0.20$0.01$14.00
    Current-generation SKUs only. $0 placeholders, Gemma, GPT-OSS, and models older than 18 months are left out of this comparison so GPT, Claude, and Gemini are compared like for like.
    ProviderCheapest current modelReleasedInput /MOutput /M
    OpenAI API pricingGPT-5 nano2025-08-07$0.05$0.40
    Claude API pricingClaude Haiku 4.52025-10-15$1.00$5.00
    Gemini API pricingGemini 2.5 Flash-Lite2025-06-17$0.10$0.40
    DeepSeek API pricingDeepSeek-V4-Flash-04232026-04-23$0.10$0.20

    Is Gemini cheaper than ChatGPT?

    On the API, yes or no depends on the SKU. Gemini vs ChatGPT is OpenAI token rates versus Google token rates. As of August 22, 2026, Gemini 2.5 Flash-Lite is $0.10/M input and $0.40/M output and GPT-5 nano is $0.05/M input and $0.40/M output. ChatGPT Plus is a subscription and is not API pricing.

    OpenAI vs Anthropic: which API should you use?

    OpenAI vs Anthropic is a per-model choice. Floor prices: GPT-5 nano at $0.05/M input and $0.40/M output versus Claude Haiku 4.5 at $1.00/M input and $5.00/M output. For coding, rank Claude vs GPT on the coding leaderboard, then pick the cheaper model if the SWE-bench gap is small.

    Claude vs ChatGPT: which is better?

    Claude vs ChatGPT is a per-model choice. ChatGPT API pricing starts at $0.05/M input and $0.40/M output (GPT-5 nano). Claude API pricing starts at $1.00/M input and $5.00/M output (Claude Haiku 4.5). Rank coding on SWE-bench, then pick the cheaper SKU if the score gap is small. ChatGPT Plus is not API pricing.

    What is the best LLM for coding?

    As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek is #1 for coding at 96.4%. Ranked by the SWE-bench Verified score This board also tracks SWE-bench Verified, LiveCodeBench, SciCode, Arena Code Elo. Next on the same board: GPT-5.6 Sol and Claude Opus 5. Related leaders: DeepSeek-V4-Pro-Max on LiveCodeBench at 93.5%; Claude Fable 5 on SciCode at 60.2%. Open the coding board for the full ranking next to API price.

    What is the best ChatGPT alternative?

    The best ChatGPT alternative for API work is usually Claude for quality or Gemini for cheaper volume. Floor rates as of August 22, 2026: Claude Claude Haiku 4.5 at $1.00/M input and $5.00/M output; Gemini Gemini 2.5 Flash-Lite at $0.10/M input and $0.40/M output. DeepSeek is the typical budget OpenAI alternative. Compare on this table, then rank coding separately.

    How does Azure OpenAI pricing compare?

    Azure OpenAI pricing is Microsoft-hosted GPT, often close to OpenAI list with contract discounts. This index tracks OpenAI’s public API, which starts at $0.05/M input and $0.40/M output (GPT-5 nano) as of August 22, 2026. Match the model name on Azure’s regional sheet before you migrate.

    Is there a free LLM API?

    A free LLM API is usually a trial, playground, or tight rate limit, not unlimited production tokens. Open-source weights can be free to download. Hosted Llama, DeepSeek, and Groq APIs still bill tokens. This table is the paid production index so you can see real $/M before you ship.

    How do I calculate LLM API cost?

    1. Count input tokens in the prompt and retrieved context.
    2. Estimate output tokens in the reply.
    3. Multiply each by the model’s input and output $/M rate.
    4. Add them. Cached input and batch APIs can discount part of the bill.

    LLM leaderboards

    Rank models by SWE-bench, Humanity's Last Exam, Terminal-Bench, Arena Elo, OSWorld, and other public evals. Every board keeps API prices next to the score.

    Individual LLM benchmarks

    One ranking page per eval: SWE-bench Verified, GPQA Diamond, Humanity's Last Exam, Terminal-Bench, LiveCodeBench, ARC-AGI, and Chatbot Arena. All benchmark pages.

    LLM comparison pages

    Open any two models to compare token pricing, context window, and benchmarks. Popular starting points:

    LLM pricing FAQ

    How much does the OpenAI API cost?

    OpenAI API pricing is metered per million tokens. As of August 22, 2026, the lowest OpenAI rate in this index is GPT-5 nano at $0.05/M input and $0.40/M output. GPT-class models cost more as reasoning, context, and output quality go up. Sort the table by OpenAI to see every live GPT rate. Source: https://developers.openai.com/api/docs/pricing

    How much does OpenAI pricing cost for GPT models?

    OpenAI pricing for GPT models is $0.05/M input and $0.40/M output at the low end (GPT-5 nano), then steps up for larger GPT and o-series models. Input tokens (your prompt) and output tokens (the reply) are billed separately, so long answers cost more than short ones.

    How much does the Claude API cost?

    Claude API pricing from Anthropic is per million tokens. As of August 22, 2026, the lowest Claude rate here is Claude Haiku 4.5 at $1.00/M input and $5.00/M output. Haiku is the cheap tier. Sonnet and Opus cost more for coding and long-context reasoning. Source: https://platform.claude.com/docs/en/about-claude/pricing

    How much is Anthropic API pricing versus OpenAI?

    Anthropic API pricing starts at $1.00/M input and $5.00/M output (Claude Haiku 4.5). OpenAI API pricing starts at $0.05/M input and $0.40/M output (GPT-5 nano). Compare blended 1M-in + 1M-out cost in the table, then open a Claude vs GPT comparison page for benchmarks.

    How much does the Gemini API cost?

    Gemini API pricing from Google is per million tokens. The lowest Gemini rate in this index is Gemini 2.5 Flash-Lite at $0.10/M input and $0.40/M output. Flash variants are built for volume. Pro and thinking variants cost more. Source: https://ai.google.dev/gemini-api/docs/pricing

    Is Gemini cheaper than ChatGPT?

    On the API, Gemini vs ChatGPT is a token-price comparison: ChatGPT uses OpenAI rates. As of August 22, 2026, the cheapest Gemini in this index is Gemini 2.5 Flash-Lite at $0.10/M input and $0.40/M output, versus GPT-5 nano at $0.05/M input and $0.40/M output. ChatGPT Plus is a flat subscription and is not the same as API billing.

    What is the cheapest LLM right now?

    As of August 22, 2026, Gemma 3 4B by Google is the cheapest LLM in this index at $0.02/M input and $0.04/M output. That is the lowest blended hosted API rate we track, not a self-host electricity cost. Sort by total cost to see the full cheapest-LLM ranking.

    What is the cheapest OpenAI model?

    GPT-5 nano is the cheapest OpenAI model in this index at $0.05/M input and $0.40/M output. It is the usual starting point for high-volume classification, routing, and cheap drafts before you pay for a larger GPT.

    OpenAI vs Anthropic: which API is cheaper?

    OpenAI vs Anthropic depends on the model, not the brand. Floor prices as of August 22, 2026: OpenAI GPT-5 nano at $0.05/M input and $0.40/M output; Anthropic Claude Haiku 4.5 at $1.00/M input and $5.00/M output. For production, compare the specific GPT and Claude SKUs you would actually call, plus SWE-bench if you are coding.

    Claude vs GPT: which is better for coding?

    Claude vs GPT for coding is not decided by price alone. Rank the coding leaderboard by SWE-bench and Terminal-Bench, then check API price on the same row. Sonnet-class Claude models and GPT-class OpenAI models trade the lead as evals refresh. Use the coding board, not a single blog screenshot.

    What is the best LLM for coding?

    As of August 22, 2026, DeepSeek-V4-Pro-0813 by DeepSeek is #1 for coding at 96.4%. Ranked by the SWE-bench Verified score This board also tracks SWE-bench Verified, LiveCodeBench, SciCode, Arena Code Elo. Next on the same board: GPT-5.6 Sol and Claude Opus 5. Related leaders: DeepSeek-V4-Pro-Max on LiveCodeBench at 93.5%; Claude Fable 5 on SciCode at 60.2%.

    What is the best LLM overall?

    “Best LLM” depends on the job: chat quality, coding, agents, or cost. This page’s LLM leaderboard ranks 282+ models on public evals next to live API prices. Use Overall for a general ranking, Coding for SWE-bench, and Pricing when token cost is the constraint.

    What is the best open source LLM?

    The best open source LLM depends on the eval. Use the open-source leaderboard for GPQA, HLE, and SWE-bench among open weights (Llama, Qwen, DeepSeek, Kimi). The cheapest open-weight API in this index is Gemma 3 4B at $0.02/M input and $0.04/M output. Weights can be free while the hosted API still bills tokens.

    How is LLM API pricing calculated?

    LLM API pricing is almost always per million tokens. Input tokens are the prompt plus any context you send. Output tokens are the model’s reply. A blended 1M input + 1M output figure lets you compare OpenAI, Claude, Gemini, DeepSeek, and Grok on one scale. Cached input, batch APIs, and image/video models use different meters.

    What is a token in an LLM?

    A token is a chunk of text the model reads or writes, roughly 0.75 words in English. LLM APIs bill by tokens, not by characters or requests. A 1,000-word prompt is about 1,300 tokens. Output is usually more expensive per token than input, so verbose answers dominate the bill.

    How do I compare LLM API prices?

    Filter the live table by provider, sort by input, output, or blended cost, then open a model page for context window and benchmarks. This LLM comparison tracks 282+ models. For Gemini vs ChatGPT or OpenAI vs Anthropic, use a two-model comparison URL so price and evals sit on one page.

    Is DeepSeek cheaper than OpenAI?

    Usually yes on raw tokens. DeepSeek’s lowest rate here is DeepSeek-V4-Flash-0423 at $0.10/M input and $0.20/M output, versus OpenAI’s GPT-5 nano at $0.05/M input and $0.40/M output. Confirm the hosted provider in the row, then check coding evals before you switch a production stack.

    How much does the Grok API cost?

    Grok API pricing from xAI is per million tokens. The lowest Grok rate in this index is Grok-4.1 Fast Non-Reasoning at $0.20/M input and $0.50/M output. Compare Grok against Claude and GPT on the same blended 1M+1M scale in the table.

    Where is the LLM leaderboard?

    The LLM leaderboard lives on this same tool: open the Overall board to rank models by Arena, GPQA, HLE, and SWE-bench next to API price. Coding, Agentic, and Open Source boards use the same prices. The snapshot is labeled August 22, 2026.

    How often is this LLM pricing table updated?

    The table is rebuilt from published API rates and public evals. This snapshot is labeled August 22, 2026. Open a model page for the exact input, output, and blended token cost, and treat the index as current, not a one-off blog post.

    Claude vs ChatGPT: which is better?

    Claude vs ChatGPT is a job split, not a single winner. ChatGPT (OpenAI GPT models) and Claude (Anthropic) both have cheap and expensive SKUs. Rank coding on the SWE-bench board, then compare live API price on the same row. ChatGPT Plus is a chat subscription. Claude API and GPT API are token bills.

    How much does ChatGPT API pricing cost?

    ChatGPT API pricing is OpenAI token pricing. As of August 22, 2026, the lowest GPT rate in this index is GPT-5 nano at $0.05/M input and $0.40/M output. ChatGPT Plus is a flat monthly fee for the chat app and is not the same meter as the API.

    What is the best ChatGPT alternative?

    The best ChatGPT alternative depends on the job. For writing and coding, Claude (Claude Haiku 4.5 from $1.00/M input and $5.00/M output) is the usual API swap. For cheap volume, Gemini (Gemini 2.5 Flash-Lite at $0.10/M input and $0.40/M output) or DeepSeek often undercut GPT. Open the coding leaderboard if the workload is software.

    How does Azure OpenAI pricing compare to OpenAI?

    Azure OpenAI pricing is Microsoft’s billed GPT rates, often close to OpenAI list with enterprise discounts and regional SKUs. This table tracks OpenAI’s public API. As of August 22, 2026, OpenAI starts at $0.05/M input and $0.40/M output (GPT-5 nano). Confirm the Azure price list for your region before you migrate.

    Is there a free LLM API?

    Free LLM APIs are usually trials or tight rate limits. Open-source weights can be free to run yourself. This table lists paid hosted token rates so you can compare production cost.

    GPT vs Claude: which API should I use?

    GPT vs Claude: pick the SKU, not the brand. Floor prices as of August 22, 2026 are OpenAI GPT-5 nano at $0.05/M input and $0.40/M output and Anthropic Claude Haiku 4.5 at $1.00/M input and $5.00/M output. For coding, rank Claude vs GPT on SWE-bench, then take the cheaper model if the score gap is small.

    How do I calculate LLM API cost?

    To calculate LLM API cost: (1) count input tokens in the prompt, (2) estimate output tokens in the reply, (3) multiply each by the model’s $/M rate, (4) add them. Example: 100k input and 20k output at $1/M and $5/M is $0.10 + $0.10 = $0.20. Cached input and batch APIs discount some of that.

    What is the difference between input and output tokens?

    Input tokens are everything you send (system prompt, user text, retrieved context). Output tokens are the model’s reply. Output is usually several times more expensive per million. A long answer can cost more than a long prompt even when the prompt has more words.

    Does a bigger context window cost more?

    A bigger context window does not raise the sticker $/M by itself. You pay for the tokens you actually send. Filling a 1M window costs far more than a 8k prompt on the same model. Some labs also add long-context surcharges above a threshold. Check the model page for that.

    Llama vs GPT: is open source cheaper?

    Llama vs GPT: hosted Llama APIs can undercut GPT, but self-hosting still costs GPUs. As of August 22, 2026, the cheapest Meta/Llama rate here is Llama 4 Scout at $0.08/M input and $0.30/M output, versus OpenAI GPT-5 nano at $0.05/M input and $0.40/M output. Weights can be free. The API is not.

    What is an OpenAI alternative for cheaper tokens?

    A cheaper OpenAI alternative is usually DeepSeek (DeepSeek-V4-Flash-0423 at $0.10/M input and $0.20/M output) or Gemini (Gemini 2.5 Flash-Lite at $0.10/M input and $0.40/M output). Confirm coding evals before you swap a production GPT. Open-source APIs can be cheaper still if quality holds.

    ChatGPT Plus vs API: which should I pay for?

    ChatGPT Plus is a monthly chat-app subscription with a usage cap. The ChatGPT API (OpenAI) bills per token with no Plus login. Use Plus for interactive chat. Use the API when your product, agent, or batch job needs programmatic calls. They are different products with different bills.

    Which LLM has the cheapest output tokens?

    Output tokens usually dominate the bill. As of August 22, 2026, Gemma 3 4B is the cheapest blended rate in this index at $0.02/M input and $0.04/M output. Sort the table by output $/M if your app writes long answers, logs, or code.