Model catalogue
Everything we serve, with the numbers most providers leave out. Quantisation is a published field, not a footnote — you can see exactly what precision your tokens are generated at before you spend anything.
17 models serving now · 3 more coming soon
Experimental A/B of the Sports-1 grounding stack on a different base model.
Experimental sibling of Sports-1: Infersia's proprietary sports grounding stack — live betting lines, player props, scores, standings, results and roster tools, the settlement rulebook and the betslip reader — on a 27-billion-parameter dense vision model. Dense means all 27B parameters compute on every token (Sports-1's mixture-of-experts activates 18B of its 320B), which is what the higher output price pays for. Multimodal (reads betslip screenshots). Here to be compared, not to replace Sports-1. 1M-token context (long-context requests are served text-only).
- Parameters
- —
- Max context
- 977K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 977K | 7.35s | — | 99.95% | $0.75 | $3.25 | $0.10 | live |
Infersia's proprietary sports grounding stack on a 320-billion-parameter sparse mixture-of-experts base — 18B parameters active per token, which is why big-model quality prices this low. Live sports data the gateway fetches itself, mid-response: betting lines — moneyline, spread and totals — and player props (anytime and first touchdown scorer, passing, rushing and receiving yards, points, rebounds, assists, home runs, strikeouts, goals and more) from named bookmakers for every major sport (NFL, NBA, MLB, NHL, college, UFC, tennis, golf, cricket, AFL, NRL and soccer leagues worldwide), roster lookups for the US leagues, plus deep soccer coverage: fixture odds, league tables, fixtures and results, squads and transfers. Send an ordinary chat completion and grounding is automatic — no tool loop to implement, no data-provider key to hold. Prices are quoted from a named bookmaker in the format its market actually uses, never an average and never converted. Bring your own tools and they take precedence. 1M-token context window.
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 1024K | 9.45s | — | 99.96% | $0.55 | $0.85 | $0.07 | live |
DeepSeek V4.1 Flash is a sparse mixture-of-experts model and the first built on DeepSeek's Causal Encoder-Decoder architecture, activating 8B parameters on input and 16B on output. It accepts text and images, holds a 1,048,576-token context, and supports tool calling, structured output and reasoning control. The asymmetric activation makes it unusually cheap on long inputs, which suits whole-repository code work, long document synthesis and agent runs that accumulate a large history.
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 1024K | — | — | 99.78% | $0.35 | $1.25 | $0.007 | live |
DeepSeek V4 Flash is a 284B mixture-of-experts model activating roughly 13B parameters per token, with a 1M context window. It uses compressed latent attention, so a very long prompt costs far less memory than its length suggests. Built for work that needs a whole corpus in one pass — repository-wide code analysis, long document synthesis and multi-step reasoning over large inputs.
- Parameters
- 284B (13B active)
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 1024K | 617ms | 57 tok/s | 99.96% | $0.10 | $0.20 | $0.02 | live |
The GA release of DeepSeek V4 Pro: a large mixture-of-experts model with a one-million-token context window and strong tool-calling and agentic behaviour.
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 1024K | — | — | 99.96% | $1.32 | $3.96 | $0.044 | live |
Natively multimodal 276B mixture-of-experts model from Thinking Machines Lab, activating roughly 12B parameters per token. Text, image and audio in; text out. Strong on agentic and long-horizon work.
- Parameters
- 276B (12B active)
- Max context
- 512K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 512K | 851ms | — | 99.96% | $0.45 | $1.20 | $0.10 | live |
xAI's frontier model, with strong coding, knowledge-work and STEM performance and a 500k-token context. Accepts text and images. Served through a third party; weights are not published.
- Parameters
- —
- Max context
- 488K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 488K | 885ms | — | 99.97% | $2.00 | $6.00 | $0.50 | live |
Muse Glimmer is Meta's 30B dense vision-language model tuned for agent workloads: tool use, long-horizon tasks and failure recovery. It accepts text and image input, supports a 131K context window, tool calling and structured output, and carries a January 2026 knowledge cutoff. Built to run always-on agents; strong on settlement-style arithmetic and reliable tool selection.
- Parameters
- 31.8B
- Max context
- 128K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 128K | — | — | 99.96% | $0.35 | $1.50 | — | live |
Qwen3.6 35B-A3B is a mixture-of-experts model with 35.9B total parameters and roughly 3B active per token, giving the quality of a large model at the speed of a much smaller one. It accepts text and image input, supports a 262K context window, and handles tool calling and structured output. Well suited to coding agents and long-context work where a whole repository or document set has to stay in the conversation.
- Parameters
- 36B (3B active)
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 256K | 3ms | — | 99.97% | $0.12 | $0.99 | $0.05 | live |
Qwen3.8 27B is a dense hybrid model — 48 Gated DeltaNet linear-attention layers and 16 gated full-attention layers — with native text, image and video input and a 262K context. Thinking is on by default with tunable reasoning effort, and it posts the family's strongest agentic scores to date (OSWorld-Verified 84.3, AndroidWorld 81.9). Built for computer-use and long-horizon agent work on a single GPU.
- Parameters
- 27B
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 256K | 26.41s | — | 99.97% | $0.45 | $3.20 | $0.05 | live |
Qwen3 8B is a dense instruction-tuned model with a hybrid thinking mode for switching between reasoning and direct answers. It offers strong general reasoning and multilingual performance for its size, with tool calling and a 32K context window. A sensible default for agent loops and high-volume tasks that do not need a frontier model.
- Parameters
- 8.2B
- Max context
- 40K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| AWQ int4free | 128K | — | — | 64.46% | Free | Free | — | down |
| Undisclosed | 128K | — | — | 99.96% | $0.05 | $0.15 | — | live |
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
- Parameters
- —
- Max context
- 512K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 512K | — | — | 99.96% | $0.35 | $1.40 | $0.07 | live |
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 1024K | 22.30s | — | 99.95% | $0.09 | $0.30 | $0.018 | live |
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 1024K | — | — | 99.95% | $1.65 | $5.20 | $0.31 | live |
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 1024K | — | — | 99.95% | $1.20 | $3.60 | $0.24 | live |
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
- Parameters
- 295B (21B active)
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 256K | — | — | 99.95% | $0.16 | $0.68 | $0.04 | live |
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 1024K | — | — | 99.96% | $0.168 | $0.336 | $0.003 | live |
Currently unavailable for maintenance.
A 35B mixture-of-experts model activating roughly 3B parameters per token, with a 262,144-token context. Built for coding and long-horizon agentic work, and strong on SWE-Bench and terminal benchmarks for its size.
- Parameters
- 36B (3B active)
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| NVFP4 | 256K | — | — | 92.94% | currently unavailable |
Currently unavailable for maintenance.
A dense 9B model with a 262,144-token context, the lightweight member of the Ornith 1.5 family. Built for coding and agentic work at a size that runs on a single GPU.
- Parameters
- 9.4B
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| BF16 | 256K | — | — | 100.00% | currently unavailable |
Currently unavailable for maintenance.
The flagship of the Ornith 1.5 family: a 397B mixture-of-experts model activating ten experts per token, with a 262,144-token context. Built for agentic coding and long-horizon software work, scoring 86.1 on Terminal-Bench 2.1 and 86.0 on SWE-bench Verified.
- Parameters
- 397B
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| NVFP4 | 256K | 79ms | — | 100.00% | currently unavailable |
On the roadmap
Listed with the same detail as everything else, so you can see what’s coming rather than guessing. These are not yet servable and will not resolve through the API.
zeroentropy/zerank-2
zerank-2 is the reranker Notion AI ran in production before acquiring its maker. A Qwen3-4B cross-encoder that rescores retrieved documents against the query, it tops published NDCG@10 comparisons against commercial rerankers, with particular strength in finance, legal, medical and code retrieval. Released Apache-2.0.
4B · 32K context
qwen/qwen3-30b-a3b-instruct-2507
Qwen3 30B-A3B Instruct is a mixture-of-experts model with 30.5B total parameters and 3.3B active per token. It has a native 262K context window rather than one extended after training, plus tool calling and structured output, making it a strong fit for agent workloads that accumulate long histories.
30.5B (3.3B active) · 256K context
google/gemma-4-26b-a4b-it
Gemma 4 26B-A4B is a sparse mixture-of-experts model from Google with roughly 4B active parameters, accepting text, image and audio input. It pairs multimodal breadth with the throughput of a far smaller dense model, and supports a 262K context window.
26.5B (3.8B active) · 256K context
qwen/qwen3-235b-a22b
235B MoE with 22B active parameters. Frontier-adjacent quality at Q4 on two A100s, in a price band where incumbents still carry substantial margin.
235B (22B active) · 128K context
openai/gpt-oss-120b
Open-weight MoE with 5.1B active parameters. The best quality-per-dollar on the board — a single A100 serves it at MXFP4, which is where the undercut room is.
117B (5.1B active) · 128K context
meta-llama/llama-4-scout
109B MoE with 17B active parameters and a very long context window. Strong at long-document work where the context length is the product.
109B (17B active) · 320K context
qwen/qwen3-1.7b
Qwen3 1.7B is the smallest model in the Qwen3 family, built for tasks where latency matters more than depth. It supports the same hybrid thinking mode as its larger siblings and handles classification, routing, extraction and short-form generation well.
1.7B · 32K context