GLM 5.2 Pricing
Z.AI ยท GLM 5 ยท Live
GLM 5.2 costs $1.4 per million input tokens and $4.4 per million output tokens on the Z.AI API. Cached input reads are billed at $0.26 per million tokens. It supports a 1,048,576-token context window and up to 128,000 output tokens per request. The model supports extended reasoning, image input, tool use, long context. Released June 16, 2026. Pricing verified September 14, 2026.
GLM 5.2 API pricing
GLM 5.2 API pricing is $1.4 per million input tokens and $4.4 per million output tokens, verified September 14, 2026.
- Input (per 1M tokens)
- $1.4
- Output (per 1M tokens)
- $4.4
- Cached input (per 1M tokens)
- $0.26
- Context window
- 1,048,576 tokens
- Max output
- 128,000 tokens
- Released
- June 16, 2026
Capabilities
GLM 5.2 supports extended reasoning, image input, tool use, long context.
Price history
See full GLM 5.2 price historyGLM 5.2 API pricing has not changed since June 28, 2026.
Price-stable since June 28, 2026.
Frequently asked questions
Short answers to the most common questions about GLM 5.2 pricing, limits, and capabilities.
Notes
Chinese frontier open-source model from Z.AI. Standalone pay-per-token API launched June 16, 2026 at $1.40/$4.40 per million. 744B parameters, 1M context window, MIT license (fully open). Highest scoring open-weight model on artificial analysis intelligence index (51). Beats GPT-5.5 on Frontier SWE coding benchmark; trails Claude Opus 4.8 by less than 1 percentage point. Cached input at $0.26/M cuts costs ~80%. Z.AI founder told Elon Musk that open-weight Fable-tier capability will arrive before Q1 2027. [2026-08-17] Prices re-verified UNCHANGED ($1.40 input / $4.40 output / $0.26 cache read) via requesty.ai, developer.puter.com, OpenRouter, pricepertoken and aipricing.guru. CONFIG FLAG: multiple sources this run state max output is 128,000 tokens per response, vs the 32,768 stored here. Left unchanged pending a second confirming run โ no dollar impact, but if it holds, raise Max Output Tokens 32,768 โ 128,000. [2026-08-24] CONFIG FLAG RESOLVED โ second consecutive run confirms max output 128,000. Updated Max Output Tokens 32,768 โ 128,000 per developer.puter.com, requesty.ai, OpenRouter and pricepertoken GLM 5.2 pages; context window 1,048,576 also re-confirmed. Dollar prices re-verified UNCHANGED at $1.40 input / $4.40 output / $0.26 cache read. Cached-input storage still listed as limited-time free โ watch for that becoming billable.
