z-ai
z-ai/…
Every z-ai model we serve, with the numbers most providers leave out — the quantisation each endpoint actually runs at, measured first-token latency, and real uptime. Model ids are copyable: pass one straight to any OpenAI-compatible client.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 1024K | 22.30s | — | 99.95% | $0.09 | $0.30 | $0.018 | live |
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 1024K | — | — | 99.95% | $1.65 | $5.20 | $0.31 | live |
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 1024K | — | — | 99.95% | $1.20 | $3.60 | $0.24 | live |