← All models
Get an API key
Thinking Machines Lab
thinkingmachines/…
1 serving nowup to 512K context
Every Thinking Machines Lab model we serve, with the numbers most providers leave out — the quantisation each endpoint actually runs at, measured first-token latency, and real uptime. Model ids are copyable: pass one straight to any OpenAI-compatible client.
Natively multimodal 276B mixture-of-experts model from Thinking Machines Lab, activating roughly 12B parameters per token. Text, image and audio in; text out. Strong on agentic and long-horizon work.
- Parameters
- 276B (12B active)
- Max context
- 512K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| Undisclosed | 512K | 851ms | — | 99.96% | $0.45 | $1.20 | $0.10 | live |
moelong-contextmultimodalvisionaudioevaluation