Pricing
Prepaid credit, billed per token at the rates below. No subscription, no minimum monthly spend, and credit that doesn’t expire.
Free
$0
Rate-limited access to models with a published free tier. Enough to evaluate the platform properly, and it costs us cents to run.
- · 10% extra on your first top-up
- · 50k tokens/day on free-tier models
- · Full API surface
Pay as you go
From $10 top-up
Add credit, spend it at the per-token rates below. Balance is decremented per request and reconciled to the nano-dollar.
- · Every model in the catalogue
- · Higher rate limits
- · Per-key spend caps
- · Usage export
Dedicated
From $500/mo
A private pool running your model — a fine-tune, or something not in our catalogue — behind a private model id only your org can resolve.
- · Reserved capacity
- · Your weights or ours
- · Region placement
- · Flat monthly fee
Per-token rates
Quoted per million tokens, with the quantisation each price applies to. A cheaper price at a lower precision is a different product, and we think you should be able to see which is which.
| Model | Quantisation | Context | Input /M | Output /M | Cached input /M |
|---|---|---|---|---|---|
| Sports-2 infersia/sports-2 | Undisclosed | 977K | $0.75 | $3.25 | $0.10 |
| Sports-1 infersia/sports-1 | Undisclosed | 1024K | $0.55 | $0.85 | $0.07 |
| DeepSeek V4.1 Flash deepseek/deepseek-v4.1-flash | Undisclosed | 1024K | $0.35 | $1.25 | $0.007 |
| DeepSeek V4 Flash deepseek/deepseek-v4-flash-0731 | Undisclosed | 1024K | $0.10 | $0.20 | $0.02 |
| DeepSeek V4 Pro deepseek/deepseek-v4-pro-0813 | Undisclosed | 1024K | $1.32 | $3.96 | $0.044 |
| Inkling Small thinkingmachines/inkling-small | Undisclosed | 512K | $0.45 | $1.20 | $0.10 |
| Grok 4.6 x-ai/grok-4.6 | Undisclosed | 488K | $2.00 | $6.00 | $0.50 |
| Muse Glimmer 30B meta-models/muse-glimmer-30b | Undisclosed | 128K | $0.35 | $1.50 | — |
| Qwen3.6 35B A3B qwen/qwen3.6-35b-a3b | Undisclosed | 256K | $0.12 | $0.99 | $0.05 |
| Qwen3.8 27B qwen/qwen3.8-27b | Undisclosed | 256K | $0.45 | $3.20 | $0.05 |
| Ornith 1.5 35B A3B ornith-ai/ornith-1.5-35b-a3b | NVFP4 | 256K | |||
| Ornith 1.5 9B ornith-ai/ornith-1.5-9b | BF16 | 256K | |||
| Ornith 1.5 397B ornith-ai/ornith-1.5-397b | NVFP4 | 256K | |||
| Qwen3 8B qwen/qwen3-8b:free | AWQ int4 | 128K | Free | Free | — |
| Qwen3 8B qwen/qwen3-8b | Undisclosed | 128K | $0.05 | $0.15 | — |
| MiniMax M3 minimax/minimax-m3 | FP8 | 512K | $0.35 | $1.40 | $0.07 |
| GLM 5.3 Flash z-ai/glm-5.3-flash | FP8 | 1024K | $0.09 | $0.30 | $0.018 |
| GLM 5.3 z-ai/glm-5.3 | FP8 | 1024K | $1.65 | $5.20 | $0.31 |
| GLM 5.2 z-ai/glm-5.2 | FP8 | 1024K | $1.20 | $3.60 | $0.24 |
| Hy3 tencent/hy3 | Undisclosed | 256K | $0.16 | $0.68 | $0.04 |
| MiMo-V2.5 xiaomi/mimo-v2.5 | FP8 | 1024K | $0.168 | $0.336 | $0.003 |
How billing actually works
Metered from the engine, not estimated
Token counts come from the inference engine’s own usage report, not from a tokeniser we run separately. Every request records whether its counts were exact, and you can see that flag on each row of your activity log.
Cached tokens are discounted, not double-billed
Cached prompt tokens are a subset of your prompt tokens, and we bill them once, at the cached rate. If you replay a long system prompt on every agent loop, that repeated prefix costs a quarter of the normal input price.
Cancelled requests bill for what was generated
Disconnect mid-stream and we bill the tokens the GPU actually produced — no more. The request appears in your log as cancelled with the exact count, so a short bill is explainable rather than mysterious.
Failures are free
If we can’t serve your request — no capacity, an upstream error, a timeout — it costs nothing. Those rows still appear in your activity log at $0.00, because you should be able to see our failure rate.