← All models
Qwen logo

Qwen

qwen/…

Get an API key
3 serving now3 on the roadmapup to 256K context

Every Qwen model we serve, with the numbers most providers leave out — the quantisation each endpoint actually runs at, measured first-token latency, and real uptime. Model ids are copyable: pass one straight to any OpenAI-compatible client.

qwen/qwen3.6-35b-a3b·Qwen logoQwen

Qwen3.6 35B-A3B is a mixture-of-experts model with 35.9B total parameters and roughly 3B active per token, giving the quality of a large model at the speed of a much smaller one. It accepts text and image input, supports a 262K context window, and handles tool calling and structured output. Well suited to coding agents and long-context work where a whole repository or document set has to stay in the conversation.

Parameters
36B (3B active)
Max context
256K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed256K3ms99.97%$0.12$0.99$0.05live
moemultimodalvisionlong-contextefficient
qwen/qwen3.8-27b·Qwen logoQwen

Qwen3.8 27B is a dense hybrid model — 48 Gated DeltaNet linear-attention layers and 16 gated full-attention layers — with native text, image and video input and a 262K context. Thinking is on by default with tunable reasoning effort, and it posts the family's strongest agentic scores to date (OSWorld-Verified 84.3, AndroidWorld 81.9). Built for computer-use and long-horizon agent work on a single GPU.

Parameters
27B
Max context
256K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed256K26.41s99.97%$0.45$3.20$0.05live
agentmultimodalvisionreasoninglong-context
qwen/qwen3-8b·Qwen logoQwen

Qwen3 8B is a dense instruction-tuned model with a hybrid thinking mode for switching between reasoning and direct answers. It offers strong general reasoning and multilingual performance for its size, with tool calling and a 32K context window. A sensible default for agent loops and high-volume tasks that do not need a frontier model.

Parameters
8.2B
Max context
40K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
AWQ int4free128K64.37%FreeFreedown
Undisclosed128K99.96%$0.05$0.15live
generalreasoningmultilingual

On the roadmap

Not yet servable — these will not resolve through the API.

Other labs we serve