Skip to content
Free tool · Prices checked 2 Oct 2026

LLM API Pricing Comparison 2026

Set your monthly token volume and see what 61 models with published prices would cost you. Every rate comes from the provider’s own pricing page; caching and batch can cut the real bill by half or more.

Your monthly usage

tokens
tokens
Type 400k, 2.5M or 1,000,000.

Cheapest option

$2.00 /mo

GPT-6 Luna

Cost per request

$0.0020

at 1,000 requests a month (GPT-6 Luna)

Price spread

100×

GPT-6 Luna vs GPT-6 Astra

ModelProviderInput costOutput costMonthly totalPer requestNote
GPT-6 LunaOpenAI$1.00$1.00$2.00$0.0020Batch and Flex: $0.05 / $0.25. Fast mode: $0.20 / $1.00. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%. Launched at 50% below…
Tencent Hy3Tencent$1.32$1.06$2.38$0.0024Tencent Cloud TokenHub, Singapore region. OpenRouter now lists the same $0.132 / $0.528. An archived OpenRouter page from 8 Jul 2026 showed $0.20 / $0.80 plus…
Qwen3.8-FlashAlibaba$1.50$0.94$2.44$0.0024International (Singapore) price. Explicit cache read $0.016.
Mistral Small 4Mistral$1.50$1.20$2.70$0.0027Mistral API standard tier. Batch processing is 50% off per Mistral's pricing FAQ.
Llama 4 ScoutMeta$1.70$1.32$3.02$0.0030Meta does not sell Llama through its own API. Prices are from Amazon Bedrock (us-east-1, on-demand, model ID meta.llama4-scout-17b-instruct-v1:0); Bedrock…
Llama 4 MaverickMeta$2.40$1.94$4.34$0.0043Meta does not sell Llama through its own API. Prices are from Amazon Bedrock (us-east-1, on-demand, model ID meta.llama4-maverick-17b-instruct-v1:0); Bedrock…
GPT-5.6 LunaOpenAI$2.00$2.40$4.40$0.0044Cut 80% on 30 Jul 2026 (launch price $1 in / $6 out). Batch and Flex: $0.10 / $0.60. Fast mode: $0.40 / $2.40. Cache writes 1.25x input; 30-minute minimum…
Codestral 25.08Mistral$3.00$1.80$4.80$0.0048Mistral API standard tier. Batch processing is 50% off per Mistral's pricing FAQ.
DeepSeek V4.1 FlashDeepSeek$3.00$2.40$5.40$0.0054Peak price shown. Off-peak is half: $0.15 in, $0.60 out, $0.003 cache hit. Peak hours 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, excluding Chinese…
MiniMax-M3MiniMax$3.00$2.40$5.40$0.0054Prices after a 'permanent 50% off' (list $0.60 / $2.40). Priority tier costs 1.5x standard.
Gemini 3.1 Flash-LiteGoogle$2.50$3.00$5.50$0.0055Audio input costs $0.50, or $0.05 cached. Cache storage $1.00 per 1M tokens per hour. Scheduled to shut down on 7 May 2027; Google names 3.5 Flash-Lite as the…
Gemini 3.5 Flash-LiteGoogle$3.00$5.00$8.00$0.0080One input price for text, image, video and audio. Cache storage $1.00 per 1M tokens per hour. Priority is $0.54 / $4.50.
Mistral Large 3Mistral$5.00$3.00$8.00$0.0080Mistral API standard tier. Batch processing is 50% off per Mistral's pricing FAQ. Much cheaper than the old Mistral Large 2 ($2/$6), which retired on 31 May…
Amazon Nova 2 LiteAmazon$3.00$5.00$8.00$0.0080Amazon Bedrock, global cross-region inference, Standard tier, us-east-1 ($0.33/$2.75 for geo cross-region and in-region). Batch and Flex tiers are $0.15/$1.25;…
GLM-4.6Z.ai$6.00$4.40$10.40$0.01Cached-input storage is free for a limited time. Same price as GLM-4.7.
Gemini 3 FlashGoogle$5.00$6.00$11.00$0.01Audio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Google calls it a legacy model; the deprecations page names…
Qwen3.8-27BAlibaba$5.00$6.00$11.00$0.01International (Singapore) price on Alibaba Model Studio / Qwen Cloud. Explicit cache read $0.05.
Tencent Hy4 previewTencent$8.34$5.00$13.34$0.01Tencent Cloud TokenHub, Singapore region.
Gemini 3.8 FlashGoogle$7.50$7.50$15.00$0.01Introductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to…
Gemini 3.7 FlashGoogle$7.50$7.50$15.00$0.01Introductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to…

Published prices

Per million tokens, cheapest input first, with cache, batch and free-tier terms.

ModelIn / 1MOut / 1MCached inBatchFree access
GPT-6 LunaOpenAI$0.10$0.50$0.0150% off–
Tencent Hy3Tencent$0.132$0.528$0.033––
Qwen3.8-FlashAlibaba$0.15$0.47$0.016–Trial credit
Mistral Small 4Mistral$0.15$0.60$0.01550% offFree tier
Llama 4 ScoutMeta$0.17$0.66–50% off–
GPT-5.6 LunaOpenAI$0.20$1.20$0.0250% off–
Llama 4 MaverickMeta$0.24$0.97–50% off–
Gemini 3.1 Flash-LiteGoogle$0.25$1.50$0.02550% offFree tier
Gemini 3.5 Flash-LiteGoogle$0.30$2.50$0.0350% offFree tier
DeepSeek V4.1 FlashDeepSeek$0.30$1.20$0.006––
MiniMax-M3MiniMax$0.30$1.20$0.06––
Codestral 25.08Mistral$0.30$0.90$0.0350% offFree tier
Amazon Nova 2 LiteAmazon$0.30$2.50$0.07550% off–
Gemini 3 FlashGoogle$0.50$3.00$0.0550% offFree tier
Qwen3.8-27BAlibaba$0.50$3.00$0.10–Trial credit
Mistral Large 3Mistral$0.50$1.50$0.0550% offFree tier

Context size at the listed input rate

A base-rate illustration, not a full-window bill estimate. Long-prompt surcharges can raise the cost; check each model’s pricing note.

ModelContext windowIn / 1MCost at listed rate
Tencent Hy3Tencent262K$0.132$0.035
Mistral Small 4Mistral256K$0.15$0.038
Codestral 25.08Mistral128K$0.30$0.038
GPT-6 LunaOpenAI1.05M$0.10$0.105
GLM-4.6Z.ai200K$0.60$0.120
Mistral Large 3Mistral256K$0.50$0.128
Qwen3.8-FlashAlibaba1M$0.15$0.150
Claude Haiku 4.5Anthropic200K$1.00$0.200
GPT-5.6 LunaOpenAI1.05M$0.20$0.210
o4-miniOpenAI200K$1.10$0.220
Llama 4 MaverickMeta1M$0.24$0.240
Kimi K2.6Moonshot262K$0.95$0.249
Gemini 3.1 Flash-LiteGoogle1.05M$0.25$0.262
MiniMax-M3MiniMax1M$0.30$0.300
Amazon Nova 2 LiteAmazon1M$0.30$0.300
GPT-5.4 miniOpenAI400K$0.75$0.300
Long-prompt surcharges are not applied here; see each model’s note.

What this index leaves out

The numbers above are list prices for uncached, real-time calls. Your bill can be lower.

Batch processing

Most large providers bill asynchronous batch jobs at half price. Tick the batch box above to apply it where it is offered.

Prompt caching

Repeated prompt prefixes bill at 10% of input or less. The LLM pricing calculator models it at your cache hit rate.

Long-prompt tiers

Some models charge more for the whole request once a prompt passes a threshold, such as 200K or 272K tokens. The note column names the tier.

Volume and enterprise deals

Negotiated rates and committed-use discounts are private, so they are not here.

Free tiers

Several providers offer rate-limited free access. The free models directory lists them.

Self-hosting

Open-weight models can run on your own GPUs. The self-host break-even tool finds the volume where that wins.

Price caching and retries at your own numbers in the LLM pricing calculator, or find your self-hosting break-even in the self-host calculator.

Not sure which model fits your budget?

Check cost against quality on your own prompts.

On a free call we look at your workload, the models you are weighing and where caching or routing would cut the bill. No vendor lock-in.

Questions

LLM pricing FAQ

Which LLM API is cheapest in 2026?

On list price per million input tokens: GPT-6 Luna ($0.10), Tencent Hy3 ($0.132), Qwen3.8-Flash ($0.15). Output often decides the bill, so compare both columns at your own mix in the calculator above.

How much does prompt caching save?

A cache hit bills at 10% of the input price on most providers, and at 5% or less on some newer models. With a long system prompt that repeats on every call, the input part of the bill can fall by 80 to 90%. Our LLM pricing calculator prices it at your cache hit rate.

How much does batch processing save?

Anthropic, OpenAI, Google, Mistral and Amazon bill asynchronous batch jobs at half the standard rate. It suits work that can wait up to a day: classification, evaluation and bulk summaries. Tick the batch box above to apply it.

Do any models have a free API tier?

Yes. Google’s Gemini API and Mistral’s free mode cover several current models, and many hosts serve open-weight models for free within limits. The free models directory lists each option with its limits and catch.

Why does the same text cost different amounts on different models?

Each model family splits text into tokens differently. Anthropic says its models from Opus 4.7 onward use a tokenizer that produces about 30% more tokens for the same text, so the real cost can rise even at the same rate.

Pricing accuracy. List prices as each provider publishes them, checked 2 Oct 2026. Prices change; confirm with the provider before you commit a budget. Models with no public price, and invite-only models, are left out.

No affiliate links. We earn nothing from any provider listed here.