← All models
Z

z-ai

z-ai/…

Get an API key
3 serving nowup to 1024K context

Every z-ai model we serve, with the numbers most providers leave out — the quantisation each endpoint actually runs at, measured first-token latency, and real uptime. Model ids are copyable: pass one straight to any OpenAI-compatible client.

z-ai/glm-5.3-flash·Zz-ai

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
FP81024K22.30s99.95%$0.09$0.30$0.018live
reasoningagentlong-contextcodingvisionfast
z-ai/glm-5.3·Zz-ai

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
FP81024K99.95%$1.65$5.20$0.31live
reasoningagentlong-contextcodingfrontier
z-ai/glm-5.2·Zz-ai

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
FP81024K99.95%$1.20$3.60$0.24live
reasoningagentlong-contextcodingfrontier

Other labs we serve