MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning, tool use, and multi-step task execution while maintaining low latency and deployment efficiency. The model excels in code generation, multi-file editing, compile-run-fix loops, and test-validated repair, showing strong results on SWE-Bench Verified, Multi-SWE-Bench, and Terminal-Bench. It also performs competitively in agentic evaluations such as BrowseComp and GAIA, effectively handling long-horizon planning, retrieval, and recovery from execution errors. Benchmarked by Artificial Analysis, MiniMax-M2 ranks among the top open-source models for composite intelligence, spanning mathematics, science, and instruction-following. Its small activation footprint enables fast inference, high concurrency, and improved unit economics, making it well-suited for large-scale agents, developer assistants, and reasoning-driven applications that require responsiveness and cost efficiency.
Back to Models
-$0.3/ M tokens$1.2/ M tokensRead:0.03/ M tokensWrite:0.375/ M tokens204.8K1.3s53.9tps



Providers
Route requests across multiple providers. Copy a provider slug to set your preference.
ProviderCache Hit RateContext
MiniMax
minimax
Uptime
24hoursDirect request success rate on AI Gateway and per-provider.
Throughput
24hoursP50 throughput on live AI Gateway traffic, in tokens per second (TPS).
Latency
24hoursP50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.
Activity
Token volume and request traffic to this model over time.
Token Consumption
Related Models
More models from MiniMax