Infersia
infersia/…
Every Infersia model we serve, with the numbers most providers leave out — the quantisation each endpoint actually runs at, measured first-token latency, and real uptime. Model ids are copyable: pass one straight to any OpenAI-compatible client.
Experimental A/B of the Sports-1 grounding stack on a different base model.
Experimental sibling of Sports-1: Infersia's proprietary sports grounding stack — live betting lines, player props, scores, standings, results and roster tools, the settlement rulebook and the betslip reader — on a 27-billion-parameter dense vision model. Dense means all 27B parameters compute on every token (Sports-1's mixture-of-experts activates 18B of its 320B), which is what the higher output price pays for. Multimodal (reads betslip screenshots). Here to be compared, not to replace Sports-1. 1M-token context (long-context requests are served text-only).
- Parameters
- —
- Max context
- 977K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 977K | 7.35s | — | 99.95% | $0.75 | $3.25 | $0.10 | live |
Infersia's proprietary sports grounding stack on a 320-billion-parameter sparse mixture-of-experts base — 18B parameters active per token, which is why big-model quality prices this low. Live sports data the gateway fetches itself, mid-response: betting lines — moneyline, spread and totals — and player props (anytime and first touchdown scorer, passing, rushing and receiving yards, points, rebounds, assists, home runs, strikeouts, goals and more) from named bookmakers for every major sport (NFL, NBA, MLB, NHL, college, UFC, tennis, golf, cricket, AFL, NRL and soccer leagues worldwide), roster lookups for the US leagues, plus deep soccer coverage: fixture odds, league tables, fixtures and results, squads and transfers. Send an ordinary chat completion and grounding is automatic — no tool loop to implement, no data-provider key to hold. Prices are quoted from a named bookmaker in the format its market actually uses, never an average and never converted. Bring your own tools and they take precedence. 1M-token context window.
- Parameters
- —
- Max context
- 1024K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 1024K | 9.45s | — | 99.96% | $0.55 | $0.85 | $0.07 | live |