LLM API Pricing Comparison 2026
Set your monthly token volume and see what 61 models with published prices would cost you. Every rate comes from the provider’s own pricing page; caching and batch can cut the real bill by half or more.
Your monthly usage
Cheapest option
$2.00 /mo
GPT-6 Luna
Cost per request
$0.0020
at 1,000 requests a month (GPT-6 Luna)
Price spread
100×
GPT-6 Luna vs GPT-6 Astra
Monthly cost by model, cheapest first
- GPT-6 Luna$2.00
- Tencent Hy3$2.38
- Qwen3.8-Flash$2.44
- Mistral Small 4$2.70
- Llama 4 Scout$3.02
- Llama 4 Maverick$4.34
- GPT-5.6 Luna$4.40
- Codestral 25.08$4.80
- DeepSeek V4.1 Flash$5.40
- MiniMax-M3$5.40
- Gemini 3.1 Flash-Lite$5.50
- Gemini 3.5 Flash-Lite$8.00
- Mistral Large 3$8.00
- Amazon Nova 2 Lite$8.00
- GLM-4.6$10.40
| Model | Provider | Input cost | Output cost | Monthly total | Per request | Note |
|---|---|---|---|---|---|---|
| GPT-6 Luna | OpenAI | $1.00 | $1.00 | $2.00 | $0.0020 | Batch and Flex: $0.05 / $0.25. Fast mode: $0.20 / $1.00. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%. Launched at 50% below… |
| Tencent Hy3 | Tencent | $1.32 | $1.06 | $2.38 | $0.0024 | Tencent Cloud TokenHub, Singapore region. OpenRouter now lists the same $0.132 / $0.528. An archived OpenRouter page from 8 Jul 2026 showed $0.20 / $0.80 plus… |
| Qwen3.8-Flash | Alibaba | $1.50 | $0.94 | $2.44 | $0.0024 | International (Singapore) price. Explicit cache read $0.016. |
| Mistral Small 4 | Mistral | $1.50 | $1.20 | $2.70 | $0.0027 | Mistral API standard tier. Batch processing is 50% off per Mistral's pricing FAQ. |
| Llama 4 Scout | Meta | $1.70 | $1.32 | $3.02 | $0.0030 | Meta does not sell Llama through its own API. Prices are from Amazon Bedrock (us-east-1, on-demand, model ID meta.llama4-scout-17b-instruct-v1:0); Bedrock… |
| Llama 4 Maverick | Meta | $2.40 | $1.94 | $4.34 | $0.0043 | Meta does not sell Llama through its own API. Prices are from Amazon Bedrock (us-east-1, on-demand, model ID meta.llama4-maverick-17b-instruct-v1:0); Bedrock… |
| GPT-5.6 Luna | OpenAI | $2.00 | $2.40 | $4.40 | $0.0044 | Cut 80% on 30 Jul 2026 (launch price $1 in / $6 out). Batch and Flex: $0.10 / $0.60. Fast mode: $0.40 / $2.40. Cache writes 1.25x input; 30-minute minimum… |
| Codestral 25.08 | Mistral | $3.00 | $1.80 | $4.80 | $0.0048 | Mistral API standard tier. Batch processing is 50% off per Mistral's pricing FAQ. |
| DeepSeek V4.1 Flash | DeepSeek | $3.00 | $2.40 | $5.40 | $0.0054 | Peak price shown. Off-peak is half: $0.15 in, $0.60 out, $0.003 cache hit. Peak hours 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, excluding Chinese… |
| MiniMax-M3 | MiniMax | $3.00 | $2.40 | $5.40 | $0.0054 | Prices after a 'permanent 50% off' (list $0.60 / $2.40). Priority tier costs 1.5x standard. |
| Gemini 3.1 Flash-Lite | $2.50 | $3.00 | $5.50 | $0.0055 | Audio input costs $0.50, or $0.05 cached. Cache storage $1.00 per 1M tokens per hour. Scheduled to shut down on 7 May 2027; Google names 3.5 Flash-Lite as the… | |
| Gemini 3.5 Flash-Lite | $3.00 | $5.00 | $8.00 | $0.0080 | One input price for text, image, video and audio. Cache storage $1.00 per 1M tokens per hour. Priority is $0.54 / $4.50. | |
| Mistral Large 3 | Mistral | $5.00 | $3.00 | $8.00 | $0.0080 | Mistral API standard tier. Batch processing is 50% off per Mistral's pricing FAQ. Much cheaper than the old Mistral Large 2 ($2/$6), which retired on 31 May… |
| Amazon Nova 2 Lite | Amazon | $3.00 | $5.00 | $8.00 | $0.0080 | Amazon Bedrock, global cross-region inference, Standard tier, us-east-1 ($0.33/$2.75 for geo cross-region and in-region). Batch and Flex tiers are $0.15/$1.25;… |
| GLM-4.6 | Z.ai | $6.00 | $4.40 | $10.40 | $0.01 | Cached-input storage is free for a limited time. Same price as GLM-4.7. |
| Gemini 3 Flash | $5.00 | $6.00 | $11.00 | $0.01 | Audio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Google calls it a legacy model; the deprecations page names… | |
| Qwen3.8-27B | Alibaba | $5.00 | $6.00 | $11.00 | $0.01 | International (Singapore) price on Alibaba Model Studio / Qwen Cloud. Explicit cache read $0.05. |
| Tencent Hy4 preview | Tencent | $8.34 | $5.00 | $13.34 | $0.01 | Tencent Cloud TokenHub, Singapore region. |
| Gemini 3.8 Flash | $7.50 | $7.50 | $15.00 | $0.01 | Introductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to… | |
| Gemini 3.7 Flash | $7.50 | $7.50 | $15.00 | $0.01 | Introductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to… |
Published prices
Per million tokens, cheapest input first, with cache, batch and free-tier terms.
| Model | In / 1M | Out / 1M | Cached in | Batch | Free access |
|---|---|---|---|---|---|
| GPT-6 LunaOpenAI | $0.10 | $0.50 | $0.01 | 50% off | – |
| Tencent Hy3Tencent | $0.132 | $0.528 | $0.033 | – | – |
| Qwen3.8-FlashAlibaba | $0.15 | $0.47 | $0.016 | – | Trial credit |
| Mistral Small 4Mistral | $0.15 | $0.60 | $0.015 | 50% off | Free tier |
| Llama 4 ScoutMeta | $0.17 | $0.66 | – | 50% off | – |
| GPT-5.6 LunaOpenAI | $0.20 | $1.20 | $0.02 | 50% off | – |
| Llama 4 MaverickMeta | $0.24 | $0.97 | – | 50% off | – |
| Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | $0.025 | 50% off | Free tier |
| Gemini 3.5 Flash-LiteGoogle | $0.30 | $2.50 | $0.03 | 50% off | Free tier |
| DeepSeek V4.1 FlashDeepSeek | $0.30 | $1.20 | $0.006 | – | – |
| MiniMax-M3MiniMax | $0.30 | $1.20 | $0.06 | – | – |
| Codestral 25.08Mistral | $0.30 | $0.90 | $0.03 | 50% off | Free tier |
| Amazon Nova 2 LiteAmazon | $0.30 | $2.50 | $0.075 | 50% off | – |
| Gemini 3 FlashGoogle | $0.50 | $3.00 | $0.05 | 50% off | Free tier |
| Qwen3.8-27BAlibaba | $0.50 | $3.00 | $0.10 | – | Trial credit |
| Mistral Large 3Mistral | $0.50 | $1.50 | $0.05 | 50% off | Free tier |
Context size at the listed input rate
A base-rate illustration, not a full-window bill estimate. Long-prompt surcharges can raise the cost; check each model’s pricing note.
| Model | Context window | In / 1M | Cost at listed rate |
|---|---|---|---|
| Tencent Hy3Tencent | 262K | $0.132 | $0.035 |
| Mistral Small 4Mistral | 256K | $0.15 | $0.038 |
| Codestral 25.08Mistral | 128K | $0.30 | $0.038 |
| GPT-6 LunaOpenAI | 1.05M | $0.10 | $0.105 |
| GLM-4.6Z.ai | 200K | $0.60 | $0.120 |
| Mistral Large 3Mistral | 256K | $0.50 | $0.128 |
| Qwen3.8-FlashAlibaba | 1M | $0.15 | $0.150 |
| Claude Haiku 4.5Anthropic | 200K | $1.00 | $0.200 |
| GPT-5.6 LunaOpenAI | 1.05M | $0.20 | $0.210 |
| o4-miniOpenAI | 200K | $1.10 | $0.220 |
| Llama 4 MaverickMeta | 1M | $0.24 | $0.240 |
| Kimi K2.6Moonshot | 262K | $0.95 | $0.249 |
| Gemini 3.1 Flash-LiteGoogle | 1.05M | $0.25 | $0.262 |
| MiniMax-M3MiniMax | 1M | $0.30 | $0.300 |
| Amazon Nova 2 LiteAmazon | 1M | $0.30 | $0.300 |
| GPT-5.4 miniOpenAI | 400K | $0.75 | $0.300 |
What this index leaves out
The numbers above are list prices for uncached, real-time calls. Your bill can be lower.
Batch processing
Most large providers bill asynchronous batch jobs at half price. Tick the batch box above to apply it where it is offered.
Prompt caching
Repeated prompt prefixes bill at 10% of input or less. The LLM pricing calculator models it at your cache hit rate.
Long-prompt tiers
Some models charge more for the whole request once a prompt passes a threshold, such as 200K or 272K tokens. The note column names the tier.
Volume and enterprise deals
Negotiated rates and committed-use discounts are private, so they are not here.
Free tiers
Several providers offer rate-limited free access. The free models directory lists them.
Self-hosting
Open-weight models can run on your own GPUs. The self-host break-even tool finds the volume where that wins.
Price caching and retries at your own numbers in the LLM pricing calculator, or find your self-hosting break-even in the self-host calculator.
Not sure which model fits your budget?
Check cost against quality on your own prompts.
On a free call we look at your workload, the models you are weighing and where caching or routing would cut the bill. No vendor lock-in.
Questions
LLM pricing FAQ
Which LLM API is cheapest in 2026?
On list price per million input tokens: GPT-6 Luna ($0.10), Tencent Hy3 ($0.132), Qwen3.8-Flash ($0.15). Output often decides the bill, so compare both columns at your own mix in the calculator above.
How much does prompt caching save?
A cache hit bills at 10% of the input price on most providers, and at 5% or less on some newer models. With a long system prompt that repeats on every call, the input part of the bill can fall by 80 to 90%. Our LLM pricing calculator prices it at your cache hit rate.
How much does batch processing save?
Anthropic, OpenAI, Google, Mistral and Amazon bill asynchronous batch jobs at half the standard rate. It suits work that can wait up to a day: classification, evaluation and bulk summaries. Tick the batch box above to apply it.
Do any models have a free API tier?
Yes. Google’s Gemini API and Mistral’s free mode cover several current models, and many hosts serve open-weight models for free within limits. The free models directory lists each option with its limits and catch.
Why does the same text cost different amounts on different models?
Each model family splits text into tokens differently. Anthropic says its models from Opus 4.7 onward use a tokenizer that produces about 30% more tokens for the same text, so the real cost can rise even at the same rate.
Pricing accuracy. List prices as each provider publishes them, checked 2 Oct 2026. Prices change; confirm with the provider before you commit a budget. Models with no public price, and invite-only models, are left out.
No affiliate links. We earn nothing from any provider listed here.