← All models

Qwen3 1.7B

Apache 2.0

qwen/qwen3-1.7b

Compact 1.7B instruct model from Alibaba. Fast and inexpensive to serve, well suited to classification, routing and extraction where a larger model is wasted.

Get an API key

Endpoints

Each row is one provider serving these weights at one precision. Prices are per million tokens.

ProviderQuantisationHardwareContextMax outInputOutputCached in
InfersiaBF16RTX A40008K4K$0.02$0.06$0.0050

Prefix caching is enabled and the discount is passed through. If your agent replays the same system prompt each iteration, the repeated portion bills at the cached rate.

What BF16 means

Full bfloat16 precision — no quantisation loss.

We serve this on vllm latest, and the quantisation is returned on every request in the x-infersia-quantization response header — so you can assert on it in your own tests rather than trusting this page.

Call it

curlbash
curl https://api.infersia.com/v1/chat/completions \
  -H "Authorization: Bearer $INFERSIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3-1.7b",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'

Pin a provider with qwen/qwen3-1.7b@infersia, or take the cheapest with the bare id. Add :free for the rate-limited free tier where one exists.