← All models
Qwen logo

Qwen3.6 35B A3B

qwen/qwen3.6-35b-a3b

Qwen3.6 35B-A3B is a mixture-of-experts model with 35.9B total parameters and roughly 3B active per token, giving the quality of a large model at the speed of a much smaller one. It accepts text and image input, supports a 262K context window, and handles tool calling and structured output. Well suited to coding agents and long-context work where a whole repository or document set has to stay in the conversation.

Get an API key

Endpoints

Each row is one provider serving these weights at one precision. Prices are per million tokens.

QuantisationContextMax outInputOutputCached in
Undisclosed256K32K$0.12$0.99$0.05

Agent config

Paste this into your coding agent and it can configure itself. The timeout is derived from this variant's measured throughput, not a guess — a long generation that outlives a client's default ceiling is the most common way a working request looks broken.

# Infersia — qwen/qwen3.6-35b-a3b
# Paste this to your agent. Values are measured, not aspirational.

provider:
  type: openai-compatible
  base_url: https://api.infersia.com/v1
  api_key: ${INFERSIA_API_KEY}   # from https://infersia.com/dashboard/keys
  model: qwen/qwen3.6-35b-a3b

limits:
  max_input_tokens: 262144      # hard limit — over this returns HTTP 413
  max_output_tokens: 32768
  timeout_seconds: 600                  # no throughput measured yet; 600 is a safe default

features:
  streaming: true
  tools: false
  vision: false
  reasoning: false
  prompt_caching: false

pricing_usd_per_million_tokens:
  input: 0.12
  output: 0.99
  cached_input: 0.05

What Undisclosed means

This model is served on capacity we buy rather than hardware we operate, so we cannot verify the precision it is served at. We would rather say that than print a figure we are guessing at.

Call it

curlbash
curl https://api.infersia.com/v1/chat/completions \
  -H "Authorization: Bearer $INFERSIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.6-35b-a3b",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'

Pin a provider with qwen/qwen3.6-35b-a3b@infersia, or take the cheapest with the bare id. Add :free for the rate-limited free tier where one exists.