Z.ai · AI model API pricing
How much does GLM-5.2 cost?
Zhipu's newest flagship — frontier-class ambitions with output tokens cheaper than most Western budget tiers.
Rate card
| Input tokens | $1.4 / 1M |
| Cached input | $0.26 / 1M |
| Output tokens | $4.4 / 1M |
| Context window | — |
What that means per month
| Workload | Tokens / month | Cost / month |
|---|---|---|
| Prototype | 1M in · 0.3M out | $2.72/mo |
| Small app | 10M in · 3M out | $27.2/mo |
| Production app | 50M in · 15M out | $136/mo |
| High volume | 200M in · 60M out | $544/mo |
| Heavy platform | 1B in · 300M out | $2,720/mo |
Standard on-demand rates from Z.ai's official pricing docs, last checked 2026-07-28. Monthly figures exclude caching and batch discounts. Confirm live pricing before committing.
Where GLM-5.2 sits on price
At the reference month of 10M input and 3M output tokens, GLM-5.2 costs $27.2, which makes it the 17th cheapest of the 30 models tracked here. That is 16× the price of Qwen-Flash, the cheapest model at this mix ($1.70). Among frontier-class models it is the cheapest of 10.
Cache economics
Cached input on GLM-5.2 bills at $0.26 per 1M — 19% of its own input rate, against a 10% median across the models here that publish a cached-read price. 24 models discount cached reads more steeply than this. On the reference month, serving 80% of that input from cache takes the input side of the bill from $14 to $4.88, leaving output untouched at $13.2 — which is why a high reuse rate changes the ranking above rather than just shaving the total.
Cache reads only. Writes are billed separately by some providers and are excluded here — see the rate note above for Z.ai's terms. Try it at your own hit rate on the AI cost calculator.
Cheaper alternatives
At a reference workload of 10M input / 3M output tokens a month, these cost less than GLM-5.2 ($27.2/mo):
-
DeepSeek V4 Pro $25.1/mo -
Claude Haiku 4.5 $25/mo -
Kimi K2.6 $21.5/mo