← All models
M
Get an API keyminimax
minimax/…
1 serving nowup to 512K context
Every minimax model we serve, with the numbers most providers leave out — the quantisation each endpoint actually runs at, measured first-token latency, and real uptime. Model ids are copyable: pass one straight to any OpenAI-compatible client.
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
- Parameters
- —
- Max context
- 512K
| Quantisation | Context | p50 TTFT | Throughput | Uptime | Input /M | Output /M | Cached in /M | Status |
|---|---|---|---|---|---|---|---|---|
| FP8 | 512K | — | — | 99.96% | $0.35 | $1.40 | $0.07 | live |
reasoningagentlong-contextvisioncoding