Ornith
ornith-ai/…
Every Ornith model we serve, with the numbers most providers leave out — the quantisation each endpoint actually runs at, measured first-token latency, and real uptime. Model ids are copyable: pass one straight to any OpenAI-compatible client.
Currently unavailable for maintenance.
The flagship of the Ornith 1.5 family: a 397B mixture-of-experts model activating ten experts per token, with a 262,144-token context. Built for agentic coding and long-horizon software work, scoring 86.1 on Terminal-Bench 2.1 and 86.0 on SWE-bench Verified.
- Parameters
- 397B
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| NVFP4 | 256K | 79ms | — | 100.00% | currently unavailable |
Currently unavailable for maintenance.
A 35B mixture-of-experts model activating roughly 3B parameters per token, with a 262,144-token context. Built for coding and long-horizon agentic work, and strong on SWE-Bench and terminal benchmarks for its size.
- Parameters
- 36B (3B active)
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| NVFP4 | 256K | — | — | 92.94% | currently unavailable |
Currently unavailable for maintenance.
A dense 9B model with a 262,144-token context, the lightweight member of the Ornith 1.5 family. Built for coding and agentic work at a size that runs on a single GPU.
- Parameters
- 9.4B
- Max context
- 256K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| BF16 | 256K | — | — | 100.00% | currently unavailable |