MiniMax M3 is the first open-weights flagship that brings coding, agentic reasoning, million-token context, and native multimodality together in one model. It handles autonomous task decomposition, tool use, and multi-step reasoning with ease — and writes code that's meant to ship, not code that just runs. Built on MiniMax's proprietary Sparse Attention architecture, M3 natively supports million-token context windows for long-horizon agents, large codebases, and long-video understanding. Its multimodality isn't bolted on — it's trained in from day one, with text and vision deeply aligned at the semantic level.
Back to Models
95.8%$0.3-0.6/ M tokens$1.2-2.4/ M tokensRead:0.06-0.12/ M tokensWrite:-/ M tokens1M1.53s134tps



Providers
Route requests across multiple providers. Copy a provider slug to set your preference.
ProviderCache Hit RateContext
MiniMax
minimax
Uptime
24hoursDirect request success rate on AI Gateway and per-provider.
Throughput
24hoursP50 throughput on live AI Gateway traffic, in tokens per second (TPS).
Latency
24hoursP50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.
Activity
Token volume and request traffic to this model over time.
Token Consumption
Related Models
More models from MiniMax