GLM-4.6
Z.ai · Released 30 Sep 2025 · glm-4.6
$0.60
Input price / 1M tokens
#17 of 61, cheapest first
$2.20
Output price / 1M tokens
#13 of 61, cheapest first
200K
Context window
#62 of 71
35
Output tokens / second
#59 of 61
Summary
GLM-4.6 is an open-weight model from Z.ai, released on 30 Sep 2025.
At $0.60 input and $2.20 output per million tokens, it is #17 of 61 on input price, cheapest first.
Its 200K-token context window ranks #62 of 71.
Measured output speed is 35 tokens per second, #59 of 61.
Older Z.ai coding model (355B total, 32B active MoE) that raised context from 128K to 200K over GLM-4.5. Open weights under MIT (commercial use allowed). Two generations behind GLM-5.3.
Benchmarks · 0 reported
No benchmark score we could source for this model. We show scores only with a source link.
Sources
- price: docs.z.ai/guides/overview/pricing, checked 2 Oct 2026
- context 200K, max output 128K: docs.z.ai/guides/llm/glm-4.6, checked 2 Oct 2026
- release date: docs.z.ai/release-notes/new-released, checked 2 Oct 2026
- license: huggingface.co/zai-org/GLM-4.6, checked 2 Oct 2026
- parameters 355B-A32B: github.com/zai-org/GLM-4.5, checked 2 Oct 2026
- output speed: artificialanalysis.ai/models/glm-4-6, checked 2 Oct 2026
Building on GLM-4.6?
Get the architecture right before the bill arrives.
We size caching, routing and fallbacks for your workload on a free call, and tell you where a cheaper model is good enough.