Qwen3.8-Max vs GLM-5.2 vs Kimi K3
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
GLM-5.2
$2.15 / 1M tokens
Largest context
GLM-5.2Kimi K3
Tie at 1.05M tokens
Fastest output
GLM-5.2
88 tokens/s
Top on Terminal-Bench 2.1
Kimi K3
88.3%
Specifications
| Spec | Qwen3.8-MaxAlibaba | GLM-5.2Z.ai | Kimi K3Moonshot |
|---|---|---|---|
| Provider | Alibaba | Z.ai | Moonshot |
| Released | 2 Aug 2026 | 16 Jun 2026 | Jul 2026 |
| Licence | Open weight | Open weight | Open weight |
| Context window | 1M tokens | 1.05M tokens | 1.05M tokens |
| Max output | 131K tokens | 131K tokens | 1.05M tokens |
| Input / 1M | $2.00* | $1.40* | $3.00* |
| Output / 1M | $6.00 | $4.40 | $15.00 |
| Cached input / 1M | $0.25 | $0.26 | $0.30 |
| Speed | 39 tokens/s | 88 tokens/s | 34 tokens/s |
| Input types | text, image, video | text | text, image, video |
| Best for | CodingAgentsReasoningLong context | CodingAgentsLong context | CodingAgentsReasoningLong context |
* Qwen3.8-Max: International (Singapore) list price. Implicit cache $0.25; explicit cache creation $2.50, explicit cache read $0.17 (Qwen Cloud). Global deployment scope is cheaper at $1.65 / $4.951. A faster qwen3.8-max-prime mode costs about 2x. Snapshot qwen3.8-max-0902 added 2 Sep 2026.
* GLM-5.2: Same price as GLM-5.3. Cached-input storage is free for a limited time.
* Kimi K3: Cache write $3.00 (5-minute TTL, default) or $6.00 (1-hour TTL); cache hits $0.30. Prices exclude taxes.
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- Qwen3.8-Max
- Alibaba's current flagship, built on the open-weight Qwen3.8-2.4T-A95B (2.4T total, 95B active MoE). The API version adds vision input, non-thinking mode and 1M context by default. Weights license allows commercial use, but products over 100M MAU or $20M monthly revenue must show the model name, and Model-as-a-Service or AI coding/office assistant businesses with over $50M revenue in 12 months need a separate license from Qwen.
- GLM-5.2
- Z.ai flagship for long-horizon tasks with a 1M-token context, 744B total and 40B active parameters. Open weights under plain MIT with no regional limits (commercial use allowed). Superseded by GLM-5.3 on 18 Aug 2026.
- Kimi K3
- Moonshot's flagship and the first open 3T-class model: 2.8T total, 104B active parameters, native vision, 1M context. Kimi K3 License allows commercial use, but Model-as-a-Service operators with over $20M revenue in 12 months need a separate agreement, and products over 100M MAU or $20M monthly revenue must display "Kimi K3".
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.