Model catalogue

Everything we serve, with the numbers most providers leave out. Quantisation is a published field, not a footnote — you can see exactly what precision your tokens are generated at before you spend anything.

Qwen3 14BApache 2.0

qwen/qwen3-14b

Dense 14B instruct model from Alibaba with a hybrid thinking mode that can be switched per request. Competitive with much larger models on reasoning, maths and code, and the point in the Qwen3 range where quality stops being the constraint for most production work.

Parameters
14.8B
Max context
128K
ProviderQuantisationEngineHardwareContextp50 TTFTThroughputUptime 30dInput /MOutput /MStatus
InfersiaAWQ int4vllm latest
32K134ms75 tok/s100.00%$0.09$0.22down
generalmultilingualagentsreasoningcode
Qwen3 8BApache 2.0

qwen/qwen3-8b

Dense 8B instruct model from Alibaba. Strong general reasoning and multilingual performance for its size, with a hybrid thinking mode. A good default for agent loops and classification where an 8B is enough and latency matters.

Parameters
8.2B
Max context
128K
ProviderQuantisationEngineHardwareContextp50 TTFTThroughputUptime 30dInput /MOutput /MStatus
InfersiafreeAWQ int4vllm 0.9.1
mock — no GPU yetRTX 3090
8K55.48%FreeFree1 live
InfersiaAWQ int4vllm 0.9.1
mock — no GPU yetRTX 3090
32K2ms31 tok/s55.48%$0.05$0.151 live
generalmultilingualagentsreasoning
Qwen3 1.7BApache 2.0

qwen/qwen3-1.7b

Compact 1.7B instruct model from Alibaba. Fast and inexpensive to serve, well suited to classification, routing and extraction where a larger model is wasted.

Parameters
1.7B
Max context
32K
ProviderQuantisationEngineHardwareContextp50 TTFTThroughputUptime 30dInput /MOutput /MStatus
InfersiaBF16vllm latest
RTX A4000
8K49 tok/s14.39%$0.02$0.06down

On the roadmap

Listed with the same detail as everything else, so you can see what’s coming rather than guessing. These are not yet servable and will not resolve through the API.