← All models
Thinking Machines Lab logo

Thinking Machines Lab

thinkingmachines/…

Get an API key
1 serving nowup to 512K context

Every Thinking Machines Lab model we serve, with the numbers most providers leave out — the quantisation each endpoint actually runs at, measured first-token latency, and real uptime. Model ids are copyable: pass one straight to any OpenAI-compatible client.

thinkingmachines/inkling-small·Thinking Machines Lab logoThinking Machines Lab

Natively multimodal 276B mixture-of-experts model from Thinking Machines Lab, activating roughly 12B parameters per token. Text, image and audio in; text out. Strong on agentic and long-horizon work.

Parameters
276B (12B active)
Max context
512K
QuantisationContextp50 TTFTThroughputUptimeInput /MOutput /MCached in /MStatus
Undisclosed512K851ms99.96%$0.45$1.20$0.10live
moelong-contextmultimodalvisionaudioevaluation

Other labs we serve