AI Model Leaderboard 2026
Compare 72 AI models on benchmark scores, context windows, prices and speed. Every number links to its source. Anything we could not confirm shows a dash, never a zero.
72
Models tracked
15
Providers
55
With sourced benchmark scores
23
Open-weight models
Model rankings
Top models by publicly reported score on the benchmark you pick. Hover a bar for its source; click it for the model page.
Terminal-Bench 4.0 · Long jobs in a real terminal
Top 10 of 10 with a score · Self-reported = the lab’s own number
- Claude Sonnet 5.570.6%
- Claude Opus 5.566.4%
- Claude Mythos 5.160.9%
- GPT-6 Astra57.9%
- Claude Fable 5.155.8%
- Claude Opus 552.3%
- Claude Fable 542%
- Grok 4.737.6%
- Gemini 3.8 Flash19.1%
- Claude Sonnet 510.3%
Top models per task
Best in class for each kind of work, ranked by one public benchmark per task. Models with no verified score are left out, not ranked last.
Coding
Terminal-Bench 4.0
- 1 Claude Sonnet 5.570.6%
- 2 Claude Opus 5.566.4%
- 3 Claude Mythos 5.160.9%
- 4 GPT-6 Astra57.9%
- 5 Claude Fable 5.155.8%
- 6 Claude Opus 552.3%
Reasoning
GPQA Diamond
- 1 GPT-6 Astra96%
- 2 GPT-5.6 Sol94.6%
- 3 Gemini 3.1 Pro94.3%
- 4 GPT-5.593.6%
- 5 Kimi K393.5%
- 6 GPT-5.6 Terra92.9%
Agents
OSWorld 2.0
- 1 Claude Fable 5.177.9%
- 2 Claude Fable 572.9%
- 3 GPT-6 Astra72.6%
- 4 GPT-6.1 Sol71.4%
- 5 GPT-5.6 Sol62.6%
- 6 GPT-6 Sol60.5%
Vision
MMMU-Pro
- 1 Gemini 3.5 Flash83.6%
- 2 Gemini 3 Flash81.2%
- 3 Gemini 3.1 Pro80.5%
- 4 Gemma 4 31B76.9%
- 5 Gemini 3.1 Flash-Lite76.8%
- 6 Gemma 4 26B A4B73.8%
Long context
Context window, tokens
- 1 Llama 4 Scout10M
- 2 GPT-6 Astra1.05M
- 3 GPT-6.1 Sol1.05M
- 4 GPT-6 Sol1.05M
- 5 GPT-6 Luna1.05M
- 6 GPT-5.6 Sol1.05M
Budget
Blended price per 1M tokens (3 in : 1 out), lower is better
- 1 GPT-6 Luna$0.20
- 2 Qwen3.8-Flash$0.23
- 3 Tencent Hy3$0.231
- 4 Mistral Small 4$0.2625
- 5 Llama 4 Scout$0.2925
- 6 Llama 4 Maverick$0.4225
Benchmark leaderboards
The top models on each benchmark. Click a row for the full model page.
DeepSWE v1.1
17 scoredGPQA Diamond
27 scoredHumanity’s Last Exam (no tools)
23 scoredTerminal-Bench 2.1
17 scoredHumanity’s Last Exam (with tools)
16 scoredSWE-bench Pro
16 scoredSWE-bench Verified
12 scoredTerminal-Bench 4.0
10 scoredFastest and most affordable
Median output speed measured by Artificial Analysis on each provider’s own API; prices from each provider’s pricing page.
Fastest, tokens per second
| Model | Provider | Tokens/s | In $/1M |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | 360 | $0.30 | |
| Gemini 3.7 Flash | 299 | $0.75 | |
| Gemini 2.5 Flash-Lite | 297 | $0.10 | |
| Gemini 3.1 Flash-Lite | 280 | $0.25 | |
| Gemini 3.8 Flash | 249 | $0.75 | |
| GPT-5.4 mini | OpenAI | 213 | $0.75 |
| Command A+ | Cohere | 211 | – |
Cheapest input price
| Model | Provider | In $/1M | Out $/1M |
|---|---|---|---|
| GPT-6 Luna | OpenAI | $0.10 | $0.50 |
| Tencent Hy3 | Tencent | $0.132 | $0.528 |
| Qwen3.8-Flash | Alibaba | $0.15 | $0.47 |
| Mistral Small 4 | Mistral | $0.15 | $0.60 |
| Llama 4 Scout | Meta | $0.17 | $0.66 |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 |
| Llama 4 Maverick | Meta | $0.24 | $0.97 |
Model comparison
Every model we track. Sort by any column, filter by provider, licence or task, and tick up to four to compare side by side.
72 of 72 models. Tick up to 4 to compare.
| Compare | Licence | Best for | |||||||
|---|---|---|---|---|---|---|---|---|---|
| GPT-6.1 Sol | OpenAI | Proprietary | 1.05M | $2.00 | $10.00 | 64 | Sep 2026 | CodingAgents | |
| Claude Sonnet 5.5 | Anthropic | Proprietary | 1M | $2.00 | $10.00 | 139 | Sep 2026 | CodingAgents | |
| Claude Opus 5.5 | Anthropic | Proprietary | 1M | $4.00 | $20.00 | 93 | Sep 2026 | CodingAgents | |
| GPT-6 Sol | OpenAI | Proprietary | 1.05M | $2.00 | $10.00 | 92 | Sep 2026 | CodingAgents | |
| GPT-6 Luna | OpenAI | Proprietary | 1.05M | $0.10 | $0.50 | 131 | Sep 2026 | BudgetAgents | |
| Grok 4.7 | xAI | Proprietary | 500K | $2.00 | $6.00 | 79 | Sep 2026 | CodingAgents | |
| Sakana Fugu Ultra v2 | Sakana AI | Proprietary | – | $5.00 | $30.00 | – | Sep 2026 | ReasoningCoding | |
| DeepSeek V4.1 Flash | DeepSeek | Open weight | 1.05M | $0.30 | $1.20 | 209 | Sep 2026 | CodingAgents | |
| GPT-6 Astra | OpenAI | Proprietary | 1.05M | $10.00 | $50.00 | 51 | Sep 2026 | CodingAgents | |
| Gemini 3.8 FlashFree tier | Proprietary | 1.05M | $0.75 | $3.75 | 249 | Sep 2026 | CodingAgents | ||
| Muse Spark 1.3 | Meta | Proprietary | 1.05M | $1.25 | $4.25 | 152 | Sep 2026 | AgentsCoding | |
| Claude Fable 5.1 | Anthropic | Proprietary | 1M | $10.00 | $50.00 | 68 | Sep 2026 | ReasoningAgents | |
| Claude Mythos 5.1Invite only | Anthropic | Proprietary | 1M | $10.00 | $50.00 | – | Sep 2026 | Reasoningresearch | |
| Qwen3.8-Flash | Alibaba | Open weight | 1M | $0.15 | $0.47 | – | Aug 2026 | BudgetCoding | |
| GLM-5.3 | Z.ai | Open weight | 1.05M | $1.40 | $4.40 | 70 | Aug 2026 | CodingAgents | |
| Qwen3.8-27B | Alibaba | Open weight | 1M | $0.50 | $3.00 | 46 | Aug 2026 | CodingAgents | |
| Gemini 3.7 FlashFree tier | Proprietary | 1.05M | $0.75 | $3.75 | 299 | Aug 2026 | CodingAgents | ||
| Qwen3.8-Max | Alibaba | Open weight | 1M | $2.00 | $6.00 | 39 | Aug 2026 | CodingAgents | |
| Tencent Hy4 previewPreview | Tencent | Open weight | 1.05M | $0.834 | $2.50 | – | Aug 2026 | CodingAgents | |
| Claude Opus 5 | Anthropic | Proprietary | 1M | $5.00 | $25.00 | 55 | Jul 2026 | CodingAgents | |
| Gemini 3.6 FlashFree tier | Proprietary | 1.05M | $0.75 | $3.75 | 185 | Jul 2026 | CodingAgents | ||
| Gemini 3.5 Flash-LiteFree tier | Proprietary | 1.05M | $0.30 | $2.50 | 360 | Jul 2026 | BudgetAgents | ||
| GPT-5.6 Sol | OpenAI | Proprietary | 1.05M | $4.00 | $20.00 | 79 | Jul 2026 | CodingAgents | |
| GPT-5.6 Terra | OpenAI | Proprietary | 1.05M | $2.00 | $12.00 | 97 | Jul 2026 | CodingBudget | |
| GPT-5.6 Luna | OpenAI | Proprietary | 1.05M | $0.20 | $1.20 | 124 | Jul 2026 | Budget |
Context window, cost and speed
Full specs side by side, with the pricing details a rate card hides: cache prices, long-prompt tiers and promotions that end.
| Model | Context | Max output | In $/1M | Out $/1M | Cached in | Tok/s | Pricing note |
|---|---|---|---|---|---|---|---|
| Llama 4 ScoutMeta | 10M | – | $0.17 | $0.66 | – | 108 | Open weightMeta does not sell Llama through its own API. Prices are from Amazon Bedrock (us-east-1, on-demand, model ID meta.llama4-scout-17b-instruct-v1:0); Bedrock batch is $0.085 in / $0.33 out. Bedrock caps output at 8K tokens. Groq retired Llama 4 Scout on 17 Jul 2026 (free/developer tiers) and Together no longer lists it on serverless. |
| GPT-6 AstraOpenAI | 1.05M | 128K | $10.00 | $50.00 | $1.00 | 51 | ProprietaryCache writes are 1.25x input; cache reads 0.1x; 30-minute minimum cache life. Batch and Flex: $5 in / $25 out. Fast mode (formerly Priority): 2x, $20 / $100. Ultrafast (Responses API, service_tier "ultrafast", Astra only): $60 in / $6 cached / $75 cache write / $300 out. Data residency endpoints +10%. |
| GPT-6.1 SolOpenAI | 1.05M | 128K | $2.00 | $10.00 | $0.10 | 64 | ProprietaryCache reads are 0.05x input (95% off), not the usual 0.1x. Cache writes 1.25x input. Batch and Flex: $1 / $5. Fast mode: $4 / $20. OpenAI says an Ultrafast option is coming 'in the coming days'; it is not on the pricing page yet. Data residency +10%. |
| GPT-6 SolOpenAI | 1.05M | 128K | $2.00 | $10.00 | $0.20 | 92 | ProprietarySuperseded by GPT-6.1 Sol, which costs the same but has cheaper cached input ($0.10). Batch and Flex: $1 / $5. Fast mode: $4 / $20. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%. Launched at 50% below GPT-5.6 Sol's promotional price. |
| GPT-6 LunaOpenAI | 1.05M | 128K | $0.10 | $0.50 | $0.01 | 131 | ProprietaryBatch and Flex: $0.05 / $0.25. Fast mode: $0.20 / $1.00. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%. Launched at 50% below GPT-5.6 Luna's price. |
| GPT-5.6 SolOpenAI | 1.05M | 128K | $4.00 | $20.00 | $0.40 | 79 | ProprietaryPromotional price since 21 Aug 2026 (launch price was $5 in / $30 out). OpenAI says it is 'available at least through November 21, 2026'. Batch and Flex: $2 / $10. Fast mode: $8 / $40. Cache writes 1.25x input; 30-minute minimum cache life. The alias gpt-5.6 routes to this model. Data residency +10%. |
| GPT-5.6 TerraOpenAI | 1.05M | 128K | $2.00 | $12.00 | $0.20 | 97 | ProprietaryCut 20% on 30 Jul 2026 (launch price $2.50 in / $15 out). Batch and Flex: $1 / $6. Fast mode: $4 / $24. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%. |
| GPT-5.6 LunaOpenAI | 1.05M | 128K | $0.20 | $1.20 | $0.02 | 124 | ProprietaryCut 80% on 30 Jul 2026 (launch price $1 in / $6 out). Batch and Flex: $0.10 / $0.60. Fast mode: $0.40 / $2.40. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%. |
| GPT-5.5OpenAI | 1.05M | 128K | $5.00 | $30.00 | $0.50 | 86 | ProprietaryThere is no separate cache-write charge (pre-GPT-5.6 models). Batch and Flex: $2.50 / $15. Fast mode: 2.5x, $12.50 / $75. Snapshot gpt-5.5-2026-04-23. GPT-5.5 Pro (gpt-5.5-pro) costs $30 / $180. Data residency +10%. |
| GPT-5.4OpenAI | 1.05M | 128K | $2.50 | $15.00 | $0.25 | 88 | ProprietaryThere is no separate cache-write charge. Batch and Flex: $1.25 / $7.50. Fast mode: $5 / $30. Snapshot gpt-5.4-2026-03-05. Data residency +10%. |
| Gemini 3.8 FlashGoogle | 1.05M | 66K | $0.75 | $3.75 | $0.075 | 249 | ProprietaryIntroductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to $1.00 in 2027. Google bills no separate cache write. One flat price up to the full 1M context. Flex is also 50% off; Priority is $1.35 / $6.75 ($2.70 / $13.50 from 2027). |
| Gemini 3.7 FlashGoogle | 1.05M | 66K | $0.75 | $3.75 | $0.075 | 299 | ProprietaryIntroductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to $1.00 in 2027. Same price as 3.8 Flash. |
| Gemini 3.6 FlashGoogle | 1.05M | 66K | $0.75 | $3.75 | $0.075 | 185 | ProprietaryLaunched at $1.50 / $7.50. It now has the same introductory price as 3.7 and 3.8 Flash through 31 Dec 2026, then $1.50 / $7.50 from 1 Jan 2027. Cache storage $0.50 per 1M tokens per hour, rising to $1.00 in 2027. |
| Gemini 3.5 FlashGoogle | 1.05M | 66K | $1.50 | $9.00 | $0.15 | 209 | ProprietaryNo promotion. It costs more than the newer 3.6, 3.7 and 3.8 Flash models. Cache storage $1.00 per 1M tokens per hour. Priority is $2.70 / $16.20. |
| Gemini 3.5 Flash-LiteGoogle | 1.05M | 66K | $0.30 | $2.50 | $0.03 | 360 | ProprietaryOne input price for text, image, video and audio. Cache storage $1.00 per 1M tokens per hour. Priority is $0.54 / $4.50. |
| Gemini 3.1 Flash-LiteGoogle | 1.05M | 66K | $0.25 | $1.50 | $0.025 | 280 | ProprietaryAudio input costs $0.50, or $0.05 cached. Cache storage $1.00 per 1M tokens per hour. Scheduled to shut down on 7 May 2027; Google names 3.5 Flash-Lite as the replacement. |
| Gemini 3.1 ProGoogle | 1.05M | 66K | $2.00 | $12.00 | $0.20 | 114 | ProprietaryPreviewStill preview. A separate endpoint, gemini-3.1-pro-preview-customtools, has the same price. Cache storage $4.50 per 1M tokens per hour. Priority is $3.60 / $21.60 up to 200K. |
| Gemini 3 FlashGoogle | 1.05M | 66K | $0.50 | $3.00 | $0.05 | 182 | ProprietaryPreviewAudio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Google calls it a legacy model; the deprecations page names gemini-3.6-flash as the replacement, with no shutdown date yet. |
| Gemini 2.5 ProGoogle | 1.05M | 66K | $1.25 | $10.00 | $0.125 | 122 | ProprietaryExisting users onlySince 18 Sep 2026 Google has limited 2.5 access to users who actively used these models before. They are not deprecated and have no shutdown date. Cache storage $4.50 per 1M tokens per hour. |
| Gemini 2.5 FlashGoogle | 1.05M | 66K | $0.30 | $2.50 | $0.03 | 191 | ProprietaryExisting users onlyAudio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Since 18 Sep 2026 access is limited to existing 2.5 users; the model is not deprecated and has no shutdown date. |
Benchmark glossary
What each benchmark measures, and why it matters for picking a model.
Terminal-Bench 4.0
Sixty-six fresh tasks done in a sandboxed terminal, from software and ML to security and science. The median task is about four hours of expert work.
tbench.aiGPQA Diamond
Expert-written biology, physics and chemistry questions built so that a web search does not give the answer away. PhD experts score around 65%.
arxiv.org/abs/2311.12022DeepSWE v1.1
Tasks written from scratch by engineers in active open-source repositories across five languages, graded by hand-written behaviour tests. Built to resist memorised answers.
deepswe.datacurve.aiHumanity’s Last Exam (no tools)
Questions written by subject experts at the edge of human knowledge, answered from the model alone. The hardest general-knowledge test in wide use.
lastexam.aiHumanity’s Last Exam (with tools)
The same expert questions, with the model allowed to search the web and run code. Closer to how a research agent works.
lastexam.aiOSWorld 2.0
Computer-use agents work through long real workflows across web apps and desktop software. We show the partial-credit score, which counts checkpoints reached; labs set the test up in slightly different ways.
os-world.github.ioSWE-bench Pro
Scale AI’s harder follow-up to SWE-bench: longer tasks in larger codebases where a fix usually touches several files. We use the public set.
scale.com/leaderboardSWE-bench Verified
Real issues from popular Python repositories, each checked by a person to be solvable. The patch counts only if the project’s own tests pass. Mostly quoted for 2025 models now.
swebench.comTerminal-Bench 2.1
The earlier Terminal-Bench task set. Scores are not comparable with 4.0, which uses new and harder tasks.
tbench.aiAutomationBench
Zapier’s test of hundreds of real business workflows, from sales and support to finance and HR, across simulated SaaS apps with messy data.
zapier.comMMMU-Pro
A harder version of MMMU: college-level questions that only make sense with the chart, diagram or photo, with distractors that punish guessing.
mmmu-benchmark.github.ioMRCR v2 (8-needle)
Eight near-identical requests are hidden in one very long conversation and the model must return the right one. A test of recall across a large context.
huggingface.co/datasets/openai/mrcrFrontierCode 1.1
Cognition’s benchmark: the agent fixes a real open-source issue and the patch is graded by hidden tests and a maintainer rubric for correctness, regressions and style.
cognition.com/blog/frontier-code-1.1ARC-AGI-2
Visual puzzles where the model must infer an unseen rule from a few examples. Built to reward generalisation over memory.
arcprize.orgAIME 2025
The 2025 American Invitational Mathematics Examination: short-answer problems that need several steps of exact reasoning.
maa.orgτ²-bench
An agent serves a simulated customer in airline, retail or telecom settings, using tools and following policy.
github.com/sierra-research/tau2-benchBrowseComp
Questions whose answers are hard to find and easy to check, so the agent has to search, read and combine many pages.
openai.com/index/browsecompOSWorld-Verified
The model operates desktop apps through screenshots, mouse and keyboard to finish everyday tasks, checked by scripts.
os-world.github.ioCharXiv Reasoning
Questions about real charts taken from arXiv papers, where the answer needs reasoning over the figure, not just reading a label.
charxiv.github.ioLiveCodeBench
Programming problems collected after a model’s training cut-off, so the model cannot have seen the answers.
livecodebench.github.ioPicking a model for a product?
We test models on your prompts, not on benchmarks.
A free call: we look at your workload, run cost-to-quality checks across providers and tell you which model to build on. No lock-in.
Questions
AI model leaderboard FAQ
Which AI model scores highest in 2026?
It depends on the work. On Terminal-Bench 4.0, the coding benchmark Anthropic, OpenAI and Google all report, the top published score is Claude Sonnet 5.5 (70.6%). On GPQA Diamond it is GPT-6 Astra (96%). Pick the benchmark closest to your work in the chart above.
What is an AI model leaderboard?
A ranking of models by public benchmark scores, price, context window and speed, so you can shortlist without reading every provider’s documentation.
How do I compare AI models side by side?
Tick up to four models in the table and press Compare, or open the Compare tab and search for them. The page link keeps your picks, so you can send it to your team.
Which AI model API is the cheapest?
By input price, GPT-6 Luna at $0.10 per million tokens. By output price, Qwen3.8-Flash at $0.47. Which is cheaper for you depends on your input to output mix; the pricing index works it out at your own volume.
Why do some cells show a dash?
We only show numbers we could trace to a source: the provider’s own pricing page or docs, the lab’s launch post, or the benchmark’s leaderboard. A dash means the number is not published, never that it is zero.
Why are some popular benchmarks missing for some models?
Labs report different benchmarks, and in 2026 Anthropic stopped reporting SWE-bench Verified and GPQA Diamond. We never mix versions of a test (Terminal-Bench 2.1 and 4.0 are different task sets), so a model only appears in a ranking for the exact test it reported.
Related
How we verify. Prices come from each provider’s pricing page and scores from the lab’s launch post, model card or the benchmark’s own leaderboard. Each model page lists its sources and the date we checked them (2 Oct 2026). Speed is measured by Artificial Analysis.
No affiliate links. HorizonLux has no commercial relationship with any provider listed here. Rankings are computed, not sold.