Skip to content
Free tool · Updated 2 Oct 2026

AI Model Leaderboard 2026

Compare 72 AI models on benchmark scores, context windows, prices and speed. Every number links to its source. Anything we could not confirm shows a dash, never a zero.

72

Models tracked

15

Providers

55

With sourced benchmark scores

23

Open-weight models

Model rankings

Top models by publicly reported score on the benchmark you pick. Hover a bar for its source; click it for the model page.

Terminal-Bench 4.0 · Long jobs in a real terminal

Top 10 of 10 with a score · Self-reported = the lab’s own number

  1. Claude Sonnet 5.570.6%
  2. Claude Opus 5.566.4%
  3. Claude Mythos 5.160.9%
  4. GPT-6 Astra57.9%
  5. Claude Fable 5.155.8%
  6. Claude Opus 552.3%
  7. Claude Fable 542%
  8. Grok 4.737.6%
  9. Gemini 3.8 Flash19.1%
  10. Claude Sonnet 510.3%

Top models per task

Best in class for each kind of work, ranked by one public benchmark per task. Models with no verified score are left out, not ranked last.

Reasoning

GPQA Diamond

  1. 1 GPT-6 Astra96%
  2. 2 GPT-5.6 Sol94.6%
  3. 3 Gemini 3.1 Pro94.3%
  4. 4 GPT-5.593.6%
  5. 5 Kimi K393.5%
  6. 6 GPT-5.6 Terra92.9%

Agents

OSWorld 2.0

  1. 1 Claude Fable 5.177.9%
  2. 2 Claude Fable 572.9%
  3. 3 GPT-6 Astra72.6%
  4. 4 GPT-6.1 Sol71.4%
  5. 5 GPT-5.6 Sol62.6%
  6. 6 GPT-6 Sol60.5%

Long context

Context window, tokens

  1. 1 Llama 4 Scout10M
  2. 2 GPT-6 Astra1.05M
  3. 3 GPT-6.1 Sol1.05M
  4. 4 GPT-6 Sol1.05M
  5. 5 GPT-6 Luna1.05M
  6. 6 GPT-5.6 Sol1.05M

Budget

Blended price per 1M tokens (3 in : 1 out), lower is better

  1. 1 GPT-6 Luna$0.20
  2. 2 Qwen3.8-Flash$0.23
  3. 3 Tencent Hy3$0.231
  4. 4 Mistral Small 4$0.2625
  5. 5 Llama 4 Scout$0.2925
  6. 6 Llama 4 Maverick$0.4225

Benchmark leaderboards

The top models on each benchmark. Click a row for the full model page.

Fastest and most affordable

Median output speed measured by Artificial Analysis on each provider’s own API; prices from each provider’s pricing page.

Fastest, tokens per second

ModelProviderTokens/sIn $/1M
Gemini 3.5 Flash-LiteGoogle360$0.30
Gemini 3.7 FlashGoogle299$0.75
Gemini 2.5 Flash-LiteGoogle297$0.10
Gemini 3.1 Flash-LiteGoogle280$0.25
Gemini 3.8 FlashGoogle249$0.75
GPT-5.4 miniOpenAI213$0.75
Command A+Cohere211–

Cheapest input price

ModelProviderIn $/1MOut $/1M
GPT-6 LunaOpenAI$0.10$0.50
Tencent Hy3Tencent$0.132$0.528
Qwen3.8-FlashAlibaba$0.15$0.47
Mistral Small 4Mistral$0.15$0.60
Llama 4 ScoutMeta$0.17$0.66
GPT-5.6 LunaOpenAI$0.20$1.20
Llama 4 MaverickMeta$0.24$0.97

Model comparison

Every model we track. Sort by any column, filter by provider, licence or task, and tick up to four to compare side by side.

72 of 72 models. Tick up to 4 to compare.

CompareLicenceBest for
GPT-6.1 SolOpenAIProprietary1.05M$2.00$10.0064Sep 2026CodingAgents
Claude Sonnet 5.5AnthropicProprietary1M$2.00$10.00139Sep 2026CodingAgents
Claude Opus 5.5AnthropicProprietary1M$4.00$20.0093Sep 2026CodingAgents
GPT-6 SolOpenAIProprietary1.05M$2.00$10.0092Sep 2026CodingAgents
GPT-6 LunaOpenAIProprietary1.05M$0.10$0.50131Sep 2026BudgetAgents
Grok 4.7xAIProprietary500K$2.00$6.0079Sep 2026CodingAgents
Sakana Fugu Ultra v2Sakana AIProprietary–$5.00$30.00–Sep 2026ReasoningCoding
DeepSeek V4.1 FlashDeepSeekOpen weight1.05M$0.30$1.20209Sep 2026CodingAgents
GPT-6 AstraOpenAIProprietary1.05M$10.00$50.0051Sep 2026CodingAgents
Gemini 3.8 FlashFree tierGoogleProprietary1.05M$0.75$3.75249Sep 2026CodingAgents
Muse Spark 1.3MetaProprietary1.05M$1.25$4.25152Sep 2026AgentsCoding
Claude Fable 5.1AnthropicProprietary1M$10.00$50.0068Sep 2026ReasoningAgents
Claude Mythos 5.1Invite onlyAnthropicProprietary1M$10.00$50.00–Sep 2026Reasoningresearch
Qwen3.8-FlashAlibabaOpen weight1M$0.15$0.47–Aug 2026BudgetCoding
GLM-5.3Z.aiOpen weight1.05M$1.40$4.4070Aug 2026CodingAgents
Qwen3.8-27BAlibabaOpen weight1M$0.50$3.0046Aug 2026CodingAgents
Gemini 3.7 FlashFree tierGoogleProprietary1.05M$0.75$3.75299Aug 2026CodingAgents
Qwen3.8-MaxAlibabaOpen weight1M$2.00$6.0039Aug 2026CodingAgents
Tencent Hy4 previewPreviewTencentOpen weight1.05M$0.834$2.50–Aug 2026CodingAgents
Claude Opus 5AnthropicProprietary1M$5.00$25.0055Jul 2026CodingAgents
Gemini 3.6 FlashFree tierGoogleProprietary1.05M$0.75$3.75185Jul 2026CodingAgents
Gemini 3.5 Flash-LiteFree tierGoogleProprietary1.05M$0.30$2.50360Jul 2026BudgetAgents
GPT-5.6 SolOpenAIProprietary1.05M$4.00$20.0079Jul 2026CodingAgents
GPT-5.6 TerraOpenAIProprietary1.05M$2.00$12.0097Jul 2026CodingBudget
GPT-5.6 LunaOpenAIProprietary1.05M$0.20$1.20124Jul 2026Budget

Context window, cost and speed

Full specs side by side, with the pricing details a rate card hides: cache prices, long-prompt tiers and promotions that end.

ModelContextMax outputIn $/1MOut $/1MCached inTok/sPricing note
Llama 4 ScoutMeta10M–$0.17$0.66–108Open weightMeta does not sell Llama through its own API. Prices are from Amazon Bedrock (us-east-1, on-demand, model ID meta.llama4-scout-17b-instruct-v1:0); Bedrock batch is $0.085 in / $0.33 out. Bedrock caps output at 8K tokens. Groq retired Llama 4 Scout on 17 Jul 2026 (free/developer tiers) and Together no longer lists it on serverless.
GPT-6 AstraOpenAI1.05M128K$10.00$50.00$1.0051ProprietaryCache writes are 1.25x input; cache reads 0.1x; 30-minute minimum cache life. Batch and Flex: $5 in / $25 out. Fast mode (formerly Priority): 2x, $20 / $100. Ultrafast (Responses API, service_tier "ultrafast", Astra only): $60 in / $6 cached / $75 cache write / $300 out. Data residency endpoints +10%.
GPT-6.1 SolOpenAI1.05M128K$2.00$10.00$0.1064ProprietaryCache reads are 0.05x input (95% off), not the usual 0.1x. Cache writes 1.25x input. Batch and Flex: $1 / $5. Fast mode: $4 / $20. OpenAI says an Ultrafast option is coming 'in the coming days'; it is not on the pricing page yet. Data residency +10%.
GPT-6 SolOpenAI1.05M128K$2.00$10.00$0.2092ProprietarySuperseded by GPT-6.1 Sol, which costs the same but has cheaper cached input ($0.10). Batch and Flex: $1 / $5. Fast mode: $4 / $20. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%. Launched at 50% below GPT-5.6 Sol's promotional price.
GPT-6 LunaOpenAI1.05M128K$0.10$0.50$0.01131ProprietaryBatch and Flex: $0.05 / $0.25. Fast mode: $0.20 / $1.00. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%. Launched at 50% below GPT-5.6 Luna's price.
GPT-5.6 SolOpenAI1.05M128K$4.00$20.00$0.4079ProprietaryPromotional price since 21 Aug 2026 (launch price was $5 in / $30 out). OpenAI says it is 'available at least through November 21, 2026'. Batch and Flex: $2 / $10. Fast mode: $8 / $40. Cache writes 1.25x input; 30-minute minimum cache life. The alias gpt-5.6 routes to this model. Data residency +10%.
GPT-5.6 TerraOpenAI1.05M128K$2.00$12.00$0.2097ProprietaryCut 20% on 30 Jul 2026 (launch price $2.50 in / $15 out). Batch and Flex: $1 / $6. Fast mode: $4 / $24. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%.
GPT-5.6 LunaOpenAI1.05M128K$0.20$1.20$0.02124ProprietaryCut 80% on 30 Jul 2026 (launch price $1 in / $6 out). Batch and Flex: $0.10 / $0.60. Fast mode: $0.40 / $2.40. Cache writes 1.25x input; 30-minute minimum cache life. Data residency +10%.
GPT-5.5OpenAI1.05M128K$5.00$30.00$0.5086ProprietaryThere is no separate cache-write charge (pre-GPT-5.6 models). Batch and Flex: $2.50 / $15. Fast mode: 2.5x, $12.50 / $75. Snapshot gpt-5.5-2026-04-23. GPT-5.5 Pro (gpt-5.5-pro) costs $30 / $180. Data residency +10%.
GPT-5.4OpenAI1.05M128K$2.50$15.00$0.2588ProprietaryThere is no separate cache-write charge. Batch and Flex: $1.25 / $7.50. Fast mode: $5 / $30. Snapshot gpt-5.4-2026-03-05. Data residency +10%.
Gemini 3.8 FlashGoogle1.05M66K$0.75$3.75$0.075249ProprietaryIntroductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to $1.00 in 2027. Google bills no separate cache write. One flat price up to the full 1M context. Flex is also 50% off; Priority is $1.35 / $6.75 ($2.70 / $13.50 from 2027).
Gemini 3.7 FlashGoogle1.05M66K$0.75$3.75$0.075299ProprietaryIntroductory price through 31 Dec 2026. From 1 Jan 2027: $1.50 input, $7.50 output, $0.15 cached input. Cache storage $0.50 per 1M tokens per hour, rising to $1.00 in 2027. Same price as 3.8 Flash.
Gemini 3.6 FlashGoogle1.05M66K$0.75$3.75$0.075185ProprietaryLaunched at $1.50 / $7.50. It now has the same introductory price as 3.7 and 3.8 Flash through 31 Dec 2026, then $1.50 / $7.50 from 1 Jan 2027. Cache storage $0.50 per 1M tokens per hour, rising to $1.00 in 2027.
Gemini 3.5 FlashGoogle1.05M66K$1.50$9.00$0.15209ProprietaryNo promotion. It costs more than the newer 3.6, 3.7 and 3.8 Flash models. Cache storage $1.00 per 1M tokens per hour. Priority is $2.70 / $16.20.
Gemini 3.5 Flash-LiteGoogle1.05M66K$0.30$2.50$0.03360ProprietaryOne input price for text, image, video and audio. Cache storage $1.00 per 1M tokens per hour. Priority is $0.54 / $4.50.
Gemini 3.1 Flash-LiteGoogle1.05M66K$0.25$1.50$0.025280ProprietaryAudio input costs $0.50, or $0.05 cached. Cache storage $1.00 per 1M tokens per hour. Scheduled to shut down on 7 May 2027; Google names 3.5 Flash-Lite as the replacement.
Gemini 3.1 ProGoogle1.05M66K$2.00$12.00$0.20114ProprietaryPreviewStill preview. A separate endpoint, gemini-3.1-pro-preview-customtools, has the same price. Cache storage $4.50 per 1M tokens per hour. Priority is $3.60 / $21.60 up to 200K.
Gemini 3 FlashGoogle1.05M66K$0.50$3.00$0.05182ProprietaryPreviewAudio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Google calls it a legacy model; the deprecations page names gemini-3.6-flash as the replacement, with no shutdown date yet.
Gemini 2.5 ProGoogle1.05M66K$1.25$10.00$0.125122ProprietaryExisting users onlySince 18 Sep 2026 Google has limited 2.5 access to users who actively used these models before. They are not deprecated and have no shutdown date. Cache storage $4.50 per 1M tokens per hour.
Gemini 2.5 FlashGoogle1.05M66K$0.30$2.50$0.03191ProprietaryExisting users onlyAudio input costs $1.00, or $0.10 cached. Cache storage $1.00 per 1M tokens per hour. Since 18 Sep 2026 access is limited to existing 2.5 users; the model is not deprecated and has no shutdown date.

Benchmark glossary

What each benchmark measures, and why it matters for picking a model.

Terminal-Bench 4.0

Sixty-six fresh tasks done in a sandboxed terminal, from software and ML to security and science. The median task is about four hours of expert work.

tbench.ai

GPQA Diamond

Expert-written biology, physics and chemistry questions built so that a web search does not give the answer away. PhD experts score around 65%.

arxiv.org/abs/2311.12022

DeepSWE v1.1

Tasks written from scratch by engineers in active open-source repositories across five languages, graded by hand-written behaviour tests. Built to resist memorised answers.

deepswe.datacurve.ai

Humanity’s Last Exam (no tools)

Questions written by subject experts at the edge of human knowledge, answered from the model alone. The hardest general-knowledge test in wide use.

lastexam.ai

Humanity’s Last Exam (with tools)

The same expert questions, with the model allowed to search the web and run code. Closer to how a research agent works.

lastexam.ai

OSWorld 2.0

Computer-use agents work through long real workflows across web apps and desktop software. We show the partial-credit score, which counts checkpoints reached; labs set the test up in slightly different ways.

os-world.github.io

SWE-bench Pro

Scale AI’s harder follow-up to SWE-bench: longer tasks in larger codebases where a fix usually touches several files. We use the public set.

scale.com/leaderboard

SWE-bench Verified

Real issues from popular Python repositories, each checked by a person to be solvable. The patch counts only if the project’s own tests pass. Mostly quoted for 2025 models now.

swebench.com

Terminal-Bench 2.1

The earlier Terminal-Bench task set. Scores are not comparable with 4.0, which uses new and harder tasks.

tbench.ai

Terminal-Bench 2.0

The 2025 Terminal-Bench task set, still quoted for older models.

tbench.ai

AutomationBench

Zapier’s test of hundreds of real business workflows, from sales and support to finance and HR, across simulated SaaS apps with messy data.

zapier.com

MMMU-Pro

A harder version of MMMU: college-level questions that only make sense with the chart, diagram or photo, with distractors that punish guessing.

mmmu-benchmark.github.io

MRCR v2 (8-needle)

Eight near-identical requests are hidden in one very long conversation and the model must return the right one. A test of recall across a large context.

huggingface.co/datasets/openai/mrcr

FrontierCode 1.1

Cognition’s benchmark: the agent fixes a real open-source issue and the patch is graded by hidden tests and a maintainer rubric for correctness, regressions and style.

cognition.com/blog/frontier-code-1.1

ARC-AGI-2

Visual puzzles where the model must infer an unseen rule from a few examples. Built to reward generalisation over memory.

arcprize.org

AIME 2025

The 2025 American Invitational Mathematics Examination: short-answer problems that need several steps of exact reasoning.

maa.org

τ²-bench

An agent serves a simulated customer in airline, retail or telecom settings, using tools and following policy.

github.com/sierra-research/tau2-bench

BrowseComp

Questions whose answers are hard to find and easy to check, so the agent has to search, read and combine many pages.

openai.com/index/browsecomp

OSWorld-Verified

The model operates desktop apps through screenshots, mouse and keyboard to finish everyday tasks, checked by scripts.

os-world.github.io

CharXiv Reasoning

Questions about real charts taken from arXiv papers, where the answer needs reasoning over the figure, not just reading a label.

charxiv.github.io

LiveCodeBench

Programming problems collected after a model’s training cut-off, so the model cannot have seen the answers.

livecodebench.github.io

Picking a model for a product?

We test models on your prompts, not on benchmarks.

A free call: we look at your workload, run cost-to-quality checks across providers and tell you which model to build on. No lock-in.

Questions

AI model leaderboard FAQ

Which AI model scores highest in 2026?

It depends on the work. On Terminal-Bench 4.0, the coding benchmark Anthropic, OpenAI and Google all report, the top published score is Claude Sonnet 5.5 (70.6%). On GPQA Diamond it is GPT-6 Astra (96%). Pick the benchmark closest to your work in the chart above.

What is an AI model leaderboard?

A ranking of models by public benchmark scores, price, context window and speed, so you can shortlist without reading every provider’s documentation.

How do I compare AI models side by side?

Tick up to four models in the table and press Compare, or open the Compare tab and search for them. The page link keeps your picks, so you can send it to your team.

Which AI model API is the cheapest?

By input price, GPT-6 Luna at $0.10 per million tokens. By output price, Qwen3.8-Flash at $0.47. Which is cheaper for you depends on your input to output mix; the pricing index works it out at your own volume.

Why do some cells show a dash?

We only show numbers we could trace to a source: the provider’s own pricing page or docs, the lab’s launch post, or the benchmark’s leaderboard. A dash means the number is not published, never that it is zero.

Why are some popular benchmarks missing for some models?

Labs report different benchmarks, and in 2026 Anthropic stopped reporting SWE-bench Verified and GPQA Diamond. We never mix versions of a test (Terminal-Bench 2.1 and 4.0 are different task sets), so a model only appears in a ranking for the exact test it reported.

Related

How we verify. Prices come from each provider’s pricing page and scores from the lab’s launch post, model card or the benchmark’s own leaderboard. Each model page lists its sources and the date we checked them (2 Oct 2026). Speed is measured by Artificial Analysis.

No affiliate links. HorizonLux has no commercial relationship with any provider listed here. Rankings are computed, not sold.