Model catalogue

Everything we serve, with the numbers most providers leave out. Quantisation is a published field, not a footnote — you can see exactly what precision your tokens are generated at before you spend anything.

17 models serving now · 3 more coming soon

infersia/sports-2·Infersia logoInfersia

Experimental A/B of the Sports-1 grounding stack on a different base model.

Experimental sibling of Sports-1: Infersia's proprietary sports grounding stack — live betting lines, player props, scores, standings, results and roster tools, the settlement rulebook and the betslip reader — on a 27-billion-parameter dense vision model. Dense means all 27B parameters compute on every token (Sports-1's mixture-of-experts activates 18B of its 320B), which is what the higher output price pays for. Multimodal (reads betslip screenshots). Here to be compared, not to replace Sports-1. 1M-token context (long-context requests are served text-only).

Parameters
Max context
977K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed977K7.35s99.95%$0.75$3.25$0.10live
sportsbettinggroundedlive-datatoolslong-context
infersia/sports-1·Infersia logoInfersia

Infersia's proprietary sports grounding stack on a 320-billion-parameter sparse mixture-of-experts base — 18B parameters active per token, which is why big-model quality prices this low. Live sports data the gateway fetches itself, mid-response: betting lines — moneyline, spread and totals — and player props (anytime and first touchdown scorer, passing, rushing and receiving yards, points, rebounds, assists, home runs, strikeouts, goals and more) from named bookmakers for every major sport (NFL, NBA, MLB, NHL, college, UFC, tennis, golf, cricket, AFL, NRL and soccer leagues worldwide), roster lookups for the US leagues, plus deep soccer coverage: fixture odds, league tables, fixtures and results, squads and transfers. Send an ordinary chat completion and grounding is automatic — no tool loop to implement, no data-provider key to hold. Prices are quoted from a named bookmaker in the format its market actually uses, never an average and never converted. Bring your own tools and they take precedence. 1M-token context window.

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed1024K9.45s99.96%$0.55$0.85$0.07live
sportsbettinggroundedlive-datatoolslong-context
deepseek/deepseek-v4.1-flash·DeepSeek logoDeepSeek

DeepSeek V4.1 Flash is a sparse mixture-of-experts model and the first built on DeepSeek's Causal Encoder-Decoder architecture, activating 8B parameters on input and 16B on output. It accepts text and images, holds a 1,048,576-token context, and supports tool calling, structured output and reasoning control. The asymmetric activation makes it unusually cheap on long inputs, which suits whole-repository code work, long document synthesis and agent runs that accumulate a large history.

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed1024K99.78%$0.35$1.25$0.007live
moelong-contextreasoningagentvision
deepseek/deepseek-v4-flash-0731·DeepSeek logoDeepSeek

DeepSeek V4 Flash is a 284B mixture-of-experts model activating roughly 13B parameters per token, with a 1M context window. It uses compressed latent attention, so a very long prompt costs far less memory than its length suggests. Built for work that needs a whole corpus in one pass — repository-wide code analysis, long document synthesis and multi-step reasoning over large inputs.

Parameters
284B (13B active)
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed1024K617ms57 tok/s99.96%$0.10$0.20$0.02live
deepseek/deepseek-v4-pro-0813·DeepSeek logoDeepSeek

The GA release of DeepSeek V4 Pro: a large mixture-of-experts model with a one-million-token context window and strong tool-calling and agentic behaviour.

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed1024K99.96%$1.32$3.96$0.044live
thinkingmachines/inkling-small·Thinking Machines Lab logoThinking Machines Lab

Natively multimodal 276B mixture-of-experts model from Thinking Machines Lab, activating roughly 12B parameters per token. Text, image and audio in; text out. Strong on agentic and long-horizon work.

Parameters
276B (12B active)
Max context
512K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed512K851ms99.96%$0.45$1.20$0.10live
moelong-contextmultimodalvisionaudioevaluation
x-ai/grok-4.6·xAI logoxAI

xAI's frontier model, with strong coding, knowledge-work and STEM performance and a 500k-token context. Accepts text and images. Served through a third party; weights are not published.

Parameters
Max context
488K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed488K885ms99.97%$2.00$6.00$0.50live
frontierlong-contextreasoningagentvisionclosed-weights
meta-models/muse-glimmer-30b·Meta logoMeta

Muse Glimmer is Meta's 30B dense vision-language model tuned for agent workloads: tool use, long-horizon tasks and failure recovery. It accepts text and image input, supports a 131K context window, tool calling and structured output, and carries a January 2026 knowledge cutoff. Built to run always-on agents; strong on settlement-style arithmetic and reliable tool selection.

Parameters
31.8B
Max context
128K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed128K99.96%$0.35$1.50live
agentmultimodalvisiontool-usereasoning
qwen/qwen3.6-35b-a3b·Qwen logoQwen

Qwen3.6 35B-A3B is a mixture-of-experts model with 35.9B total parameters and roughly 3B active per token, giving the quality of a large model at the speed of a much smaller one. It accepts text and image input, supports a 262K context window, and handles tool calling and structured output. Well suited to coding agents and long-context work where a whole repository or document set has to stay in the conversation.

Parameters
36B (3B active)
Max context
256K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed256K3ms99.97%$0.12$0.99$0.05live
moemultimodalvisionlong-contextefficient
qwen/qwen3.8-27b·Qwen logoQwen

Qwen3.8 27B is a dense hybrid model — 48 Gated DeltaNet linear-attention layers and 16 gated full-attention layers — with native text, image and video input and a 262K context. Thinking is on by default with tunable reasoning effort, and it posts the family's strongest agentic scores to date (OSWorld-Verified 84.3, AndroidWorld 81.9). Built for computer-use and long-horizon agent work on a single GPU.

Parameters
27B
Max context
256K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed256K26.41s99.97%$0.45$3.20$0.05live
agentmultimodalvisionreasoninglong-context
qwen/qwen3-8b·Qwen logoQwen

Qwen3 8B is a dense instruction-tuned model with a hybrid thinking mode for switching between reasoning and direct answers. It offers strong general reasoning and multilingual performance for its size, with tool calling and a 32K context window. A sensible default for agent loops and high-volume tasks that do not need a frontier model.

Parameters
8.2B
Max context
40K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
AWQ int4free128K64.46%FreeFreedown
Undisclosed128K99.96%$0.05$0.15live
generalreasoningmultilingual
minimax/minimax-m3·Mminimax

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Parameters
Max context
512K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
FP8512K99.96%$0.35$1.40$0.07live
reasoningagentlong-contextvisioncoding
z-ai/glm-5.3-flash·Zz-ai

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
FP81024K22.30s99.95%$0.09$0.30$0.018live
reasoningagentlong-contextcodingvisionfast
z-ai/glm-5.3·Zz-ai

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
FP81024K99.95%$1.65$5.20$0.31live
reasoningagentlong-contextcodingfrontier
z-ai/glm-5.2·Zz-ai

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
FP81024K99.95%$1.20$3.60$0.24live
reasoningagentlong-contextcodingfrontier
tencent/hy3·Ttencent

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

Parameters
295B (21B active)
Max context
256K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed256K99.95%$0.16$0.68$0.04live
moereasoningagentlong-contextcoding
xiaomi/mimo-v2.5·Xxiaomi

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

Parameters
Max context
1024K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
FP81024K99.96%$0.168$0.336$0.003live
omnimodallong-contextagentreasoningvisionaudio
Ornith logoOrnith 1.5 35B A3BCurrently unavailable
ornith-ai/ornith-1.5-35b-a3b·Ornith logoOrnith

Currently unavailable for maintenance.

A 35B mixture-of-experts model activating roughly 3B parameters per token, with a 262,144-token context. Built for coding and long-horizon agentic work, and strong on SWE-Bench and terminal benchmarks for its size.

Parameters
36B (3B active)
Max context
256K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
NVFP4256K92.94%$0.15$0.075$0.60$0.30$0.0375$0.018currently unavailable
moelong-contextreasoningagentcoding
Ornith logoOrnith 1.5 9BCurrently unavailable
ornith-ai/ornith-1.5-9b·Ornith logoOrnith

Currently unavailable for maintenance.

A dense 9B model with a 262,144-token context, the lightweight member of the Ornith 1.5 family. Built for coding and agentic work at a size that runs on a single GPU.

Parameters
9.4B
Max context
256K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
BF16256K100.00%$0.06$0.03$0.24$0.12$0.015$0.0075currently unavailable
long-contextreasoningagentcodingsmall
Ornith logoOrnith 1.5 397BCurrently unavailable
ornith-ai/ornith-1.5-397b·Ornith logoOrnith

Currently unavailable for maintenance.

The flagship of the Ornith 1.5 family: a 397B mixture-of-experts model activating ten experts per token, with a 262,144-token context. Built for agentic coding and long-horizon software work, scoring 86.1 on Terminal-Bench 2.1 and 86.0 on SWE-bench Verified.

Parameters
397B
Max context
256K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
NVFP4256K79ms100.00%$1.20$0.60$3.60$1.80$0.30$0.15currently unavailable
moelong-contextreasoningagentcodingfrontier

On the roadmap

Listed with the same detail as everything else, so you can see what’s coming rather than guessing. These are not yet servable and will not resolve through the API.

Zzerank-2

zeroentropy/zerank-2

zerank-2 is the reranker Notion AI ran in production before acquiring its maker. A Qwen3-4B cross-encoder that rescores retrieved documents against the query, it tops published NDCG@10 comparisons against commercial rerankers, with particular strength in finance, legal, medical and code retrieval. Released Apache-2.0.

4B · 32K context

Qwen logoQwen3 30B A3B Instruct 2507

qwen/qwen3-30b-a3b-instruct-2507

Qwen3 30B-A3B Instruct is a mixture-of-experts model with 30.5B total parameters and 3.3B active per token. It has a native 262K context window rather than one extended after training, plus tool calling and structured output, making it a strong fit for agent workloads that accumulate long histories.

30.5B (3.3B active) · 256K context

Google logoGemma 4 26B A4B

google/gemma-4-26b-a4b-it

Gemma 4 26B-A4B is a sparse mixture-of-experts model from Google with roughly 4B active parameters, accepting text, image and audio input. It pairs multimodal breadth with the throughput of a far smaller dense model, and supports a 262K context window.

26.5B (3.8B active) · 256K context

Qwen logoQwen3 235B A22B

qwen/qwen3-235b-a22b

235B MoE with 22B active parameters. Frontier-adjacent quality at Q4 on two A100s, in a price band where incumbents still carry substantial margin.

235B (22B active) · 128K context

OpenAI logogpt-oss 120B

openai/gpt-oss-120b

Open-weight MoE with 5.1B active parameters. The best quality-per-dollar on the board — a single A100 serves it at MXFP4, which is where the undercut room is.

117B (5.1B active) · 128K context

MLlama 4 Scout

meta-llama/llama-4-scout

109B MoE with 17B active parameters and a very long context window. Strong at long-document work where the context length is the product.

109B (17B active) · 320K context

Qwen logoQwen3 1.7B

qwen/qwen3-1.7b

Qwen3 1.7B is the smallest model in the Qwen3 family, built for tasks where latency matters more than depth. It supports the same hybrid thinking mode as its larger siblings and handles classification, routing, extraction and short-form generation well.

1.7B · 32K context