Qwen3.8-Flash vs GLM-5.2 vs Qwen3.7-Max
Pick two to four models to compare benchmark scores, prices, context windows and speed. The link updates as you pick, so you can share the exact comparison.
Cheapest blended
Qwen3.8-Flash
$0.23 / 1M tokens
Largest context
GLM-5.2
1.05M tokens
Fastest output
Qwen3.7-Max
203 tokens/s
Top on SWE-bench Pro
Qwen3.8-Flash
62.5%
Specifications
| Spec | Qwen3.8-FlashAlibaba | GLM-5.2Z.ai | Qwen3.7-MaxAlibaba |
|---|---|---|---|
| Provider | Alibaba | Z.ai | Alibaba |
| Released | 26 Aug 2026 | 16 Jun 2026 | 21 May 2026 |
| Licence | Open weight | Open weight | Proprietary |
| Context window | 1M tokens | 1.05M tokens | 1M tokens |
| Max output | 131K tokens | 131K tokens | 131K tokens |
| Input / 1M | $0.15* | $1.40* | $2.50* |
| Output / 1M | $0.47 | $4.40 | $7.50 |
| Cached input / 1M | $0.016 | $0.26 | $0.50 |
| Speed | – | 88 tokens/s | 203 tokens/s |
| Input types | text, image, video | text | text |
| Best for | BudgetCodingAgentsMultimodal | CodingAgentsLong context | CodingAgentsReasoningLong context |
* Qwen3.8-Flash: International (Singapore) price. Explicit cache read $0.016.
* GLM-5.2: Same price as GLM-5.3. Cached-input storage is free for a limited time.
* Qwen3.7-Max: International (Singapore) list price. Implicit cache $0.50; explicit cache read $0.25. Global deployment scope $1.65 / $4.951. Alibaba now lists Qwen3.7 under Legacy models; Qwen3.8-Max is cheaper ($2 / $6).
Benchmark scores
Officially reported scores only. A model with no published score on a benchmark shows a line, not a zero.
Price per 1M tokens
Context window and speed
Notes
- Qwen3.8-Flash
- Fast multimodal Qwen model based on the open-weight Qwen3.8-Flash-Next (125B total, 6B active, plus 51B n-gram embedding), an experimental preview of the Qwen4 architecture. The Qwen Community License allows commercial use with a name-display rule above 100M MAU or $20M monthly revenue, but any Model-as-a-Service or AI coding/office assistant business needs a separate license from Qwen regardless of size.
- GLM-5.2
- Z.ai flagship for long-horizon tasks with a 1M-token context, 744B total and 40B active parameters. Open weights under plain MIT with no regional limits (commercial use allowed). Superseded by GLM-5.3 on 18 Aug 2026.
- Qwen3.7-Max
- Largest model of the Qwen3.7 series, aimed at agent work, programming and long autonomous tasks. The qwen3.7-max alias points to the 2026-05-20 text-only snapshot; the 2026-06-08 snapshot adds image understanding. Not open weight.
Still torn between two?
Test them on your own prompts.
Benchmarks are someone else’s work. On a free call we look at yours and agree how to check quality, cost and latency for the models you are weighing.