Z.AI logo

GLM 5.2 Pricing

Z.AI ยท GLM 5 ยท Live

Prices verified September 14, 2026

GLM 5.2 costs $1.4 per million input tokens and $4.4 per million output tokens on the Z.AI API. Cached input reads are billed at $0.26 per million tokens. It supports a 1,048,576-token context window and up to 128,000 output tokens per request. The model supports extended reasoning, image input, tool use, long context. Released June 16, 2026. Pricing verified September 14, 2026.

GLM 5.2 API pricing

GLM 5.2 API pricing is $1.4 per million input tokens and $4.4 per million output tokens, verified September 14, 2026.

Input (per 1M tokens)
$1.4
Output (per 1M tokens)
$4.4
Cached input (per 1M tokens)
$0.26
Context window
1,048,576 tokens
Max output
128,000 tokens
Released
June 16, 2026
Estimate your GLM 5.2 cost

Capabilities

GLM 5.2 supports extended reasoning, image input, tool use, long context.

Extended reasoningImage inputTool useLong context

GLM 5.2 API pricing has not changed since June 28, 2026.

Price-stable since June 28, 2026.

Frequently asked questions

Short answers to the most common questions about GLM 5.2 pricing, limits, and capabilities.

Notes

Chinese frontier open-source model from Z.AI. Standalone pay-per-token API launched June 16, 2026 at $1.40/$4.40 per million. 744B parameters, 1M context window, MIT license (fully open). Highest scoring open-weight model on artificial analysis intelligence index (51). Beats GPT-5.5 on Frontier SWE coding benchmark; trails Claude Opus 4.8 by less than 1 percentage point. Cached input at $0.26/M cuts costs ~80%. Z.AI founder told Elon Musk that open-weight Fable-tier capability will arrive before Q1 2027. [2026-08-17] Prices re-verified UNCHANGED ($1.40 input / $4.40 output / $0.26 cache read) via requesty.ai, developer.puter.com, OpenRouter, pricepertoken and aipricing.guru. CONFIG FLAG: multiple sources this run state max output is 128,000 tokens per response, vs the 32,768 stored here. Left unchanged pending a second confirming run โ€” no dollar impact, but if it holds, raise Max Output Tokens 32,768 โ†’ 128,000. [2026-08-24] CONFIG FLAG RESOLVED โ€” second consecutive run confirms max output 128,000. Updated Max Output Tokens 32,768 โ†’ 128,000 per developer.puter.com, requesty.ai, OpenRouter and pricepertoken GLM 5.2 pages; context window 1,048,576 also re-confirmed. Dollar prices re-verified UNCHANGED at $1.40 input / $4.40 output / $0.26 cache read. Cached-input storage still listed as limited-time free โ€” watch for that becoming billable.