Model catalogue
Everything we serve, with the numbers most providers leave out. Quantisation is a published field, not a footnote — you can see exactly what precision your tokens are generated at before you spend anything.
qwen/qwen3-14b
Dense 14B instruct model from Alibaba with a hybrid thinking mode that can be switched per request. Competitive with much larger models on reasoning, maths and code, and the point in the Qwen3 range where quality stops being the constraint for most production work.
- Parameters
- 14.8B
- Max context
- 128K
| Provider | Quantisation | Engine | Hardware | Context | p50 TTFT | Throughput | Uptime 30d | Input /M | Output /M | Status |
|---|---|---|---|---|---|---|---|---|---|---|
| Infersia | AWQ int4 | vllm latest | — | 32K | 134ms | 75 tok/s | 100.00% | $0.09 | $0.22 | down |
qwen/qwen3-8b
Dense 8B instruct model from Alibaba. Strong general reasoning and multilingual performance for its size, with a hybrid thinking mode. A good default for agent loops and classification where an 8B is enough and latency matters.
- Parameters
- 8.2B
- Max context
- 128K
| Provider | Quantisation | Engine | Hardware | Context | p50 TTFT | Throughput | Uptime 30d | Input /M | Output /M | Status |
|---|---|---|---|---|---|---|---|---|---|---|
| Infersiafree | AWQ int4 | vllm 0.9.1 | mock — no GPU yetRTX 3090 | 8K | — | — | 55.48% | Free | Free | 1 live |
| Infersia | AWQ int4 | vllm 0.9.1 | mock — no GPU yetRTX 3090 | 32K | 2ms | 31 tok/s | 55.48% | $0.05 | $0.15 | 1 live |
qwen/qwen3-1.7b
Compact 1.7B instruct model from Alibaba. Fast and inexpensive to serve, well suited to classification, routing and extraction where a larger model is wasted.
- Parameters
- 1.7B
- Max context
- 32K
| Provider | Quantisation | Engine | Hardware | Context | p50 TTFT | Throughput | Uptime 30d | Input /M | Output /M | Status |
|---|---|---|---|---|---|---|---|---|---|---|
| Infersia | BF16 | vllm latest | RTX A4000 | 8K | — | 49 tok/s | 14.39% | $0.02 | $0.06 | down |
On the roadmap
Listed with the same detail as everything else, so you can see what’s coming rather than guessing. These are not yet servable and will not resolve through the API.
qwen/qwen3-235b-a22b
235B MoE with 22B active parameters. Frontier-adjacent quality at Q4 on two A100s, in a price band where incumbents still carry substantial margin.
235B (22B active) · 128K context
openai/gpt-oss-120b
Open-weight MoE with 5.1B active parameters. The best quality-per-dollar on the board — a single A100 serves it at MXFP4, which is where the undercut room is.
117B (5.1B active) · 128K context
meta-llama/llama-4-scout
109B MoE with 17B active parameters and a very long context window. Strong at long-document work where the context length is the product.
109B (17B active) · 320K context