Vendor
Z.ai models
16 models in the reference catalogue, vendor list prices per million tokens. 0 are routable on this deployment today.
Models
16
0 routable here
Median list price
$0.91
3:1 input to output per 1M tokens
Every Z.ai model
Newest first. Prices are the vendor's published rates per million tokens; click a model for the full specification.
| Model | Released | Context | Input / 1M | Output / 1M | Intelligence | Capabilities |
|---|---|---|---|---|---|---|
| GLM Flash Latest~z-ai/glm-flash-latest | Aug 27, 2026 | 1.3M | $0.07 | $0.25 | - | reasoning tools |
| GLM 5.3 Flashz-ai/glm-5.3-flash | Aug 26, 2026 | 1.3M | $0.07 | $0.25 | 46 | reasoning tools open weights |
| GLM Latest~z-ai/glm-latest | Aug 19, 2026 | 1.3M | $1.17 | $3.96 | - | reasoning tools |
| GLM 5.3z-ai/glm-5.3 | Aug 18, 2026 | 1.3M | $1.40 | $4.40 | 49 | reasoning tools open weights |
| GLM 5.2z-ai/glm-5.2 | Jun 16, 2026 | 1.0M | $0.97 | $3.04 | - | reasoning tools open weights |
| GLM 5.1z-ai/glm-5.1 | Apr 7, 2026 | 205K | $0.97 | $3.04 | - | reasoning tools open weights |
| GLM 5V Turboz-ai/glm-5v-turbo | Apr 1, 2026 | 203K | $1.20 | $4.00 | - | reasoning tools |
| GLM 5 Turboz-ai/glm-5-turbo | Mar 15, 2026 | 203K | $1.20 | $4.00 | - | reasoning tools |
| GLM 5z-ai/glm-5 | Feb 11, 2026 | 205K | $0.60 | $1.92 | - | reasoning tools open weights |
| GLM 4.7 Flashz-ai/glm-4.7-flash | Jan 19, 2026 | 203K | $0.06 | $0.40 | - | reasoning tools open weights |
| GLM 4.7z-ai/glm-4.7 | Dec 22, 2025 | 205K | $0.40 | $1.75 | - | reasoning tools open weights |
| GLM 4.6Vz-ai/glm-4.6v | Dec 8, 2025 | 131K | $0.30 | $0.90 | - | reasoning tools open weights |
| GLM 4.6z-ai/glm-4.6 | Sep 30, 2025 | 205K | $0.43 | $1.75 | - | reasoning tools open weights |
| GLM 4.5Vz-ai/glm-4.5v | Aug 11, 2025 | 66K | $0.60 | $1.80 | - | reasoning tools open weights |
| GLM 4.5z-ai/glm-4.5 | Jul 25, 2025 | 131K | $0.60 | $2.20 | - | reasoning tools open weights |
| GLM 4.5 Airz-ai/glm-4.5-air | Jul 25, 2025 | 131K | $0.13 | $0.85 | - | reasoning tools open weights |
Frequently asked
- What is the cheapest Z.ai model?
- GLM Flash Latest at $0.07 per 1M input tokens and $0.25 per 1M output tokens (vendor list price).
- Which Z.ai model has the largest context window?
- GLM Flash Latest accepts 1.3M tokens of context.
- How do I route to Z.ai models with OpenSmartRoute?
- Declare a target in targets.yaml with metadata.model set to the catalogue id (for example ~z-ai/glm-flash-latest); the router scores it against every other target on cost, quality, latency and your policies for each request.
Route Z.ai with everything else
OpenSmartRoute picks the cheapest model that meets your quality bar per request, so a Z.ai flagship handles hard prompts while small models take the rest. Price a workload or try a routing decision live.