minimax

MiniMax

Browse models from MiniMax
Models · 9
2.74Ktokens
-Cache Hit Rate

MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and video-to-video motion transfer.

The model is suited for commercial creative workflows across advertising, e-commerce, gaming, and interface design, with native audiovisual output for reference-driven generation.

Input Type
Output Type
Input$0.074-0.1185/seconds
Output$0.074-0.1185/seconds
Context-
Max Output-
796.18Mtokens
73.29%Cache Hit Rate

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning, tool use, and multi-step task execution while maintaining low latency and deployment efficiency. The model excels in code generation, multi-file editing, compile-run-fix loops, and test-validated repair, showing strong results on SWE-Bench Verified, Multi-SWE-Bench, and Terminal-Bench. It also performs competitively in agentic evaluations such as BrowseComp and GAIA, effectively handling long-horizon planning, retrieval, and recovery from execution errors. Benchmarked by Artificial Analysis, MiniMax-M2 ranks among the top open-source models for composite intelligence, spanning mathematics, science, and instruction-following. Its small activation footprint enables fast inference, high concurrency, and improved unit economics, making it well-suited for large-scale agents, developer assistants, and reasoning-driven applications that require responsiveness and cost efficiency.

Input Type
Output Type
Input$0.3/M tokens
Output$1.2/M tokens
Context-
Max Output-
1.10Btokens
82.39%Cache Hit Rate

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

Compared to its predecessor, M2.1 delivers cleaner, more concise outputs and faster perceived response times. It shows leading multilingual coding performance across major systems and application languages, achieving 49.4% on Multi-SWE-Bench and 72.5% on SWE-Bench Multilingual, and serves as a versatile agent “brain” for IDEs, coding tools, and general-purpose assistance.

To avoid degrading this model's performance, MiniMax highly recommends preserving reasoning between turns.

Input Type
Output Type
Input$0.3/M tokens
Output$1.2/M tokens
Context-
Max Output-
2.71Mtokens
0%Cache Hit Rate

M2-her text chat model, designed for role-playing, multi-turn conversations and dialogue scenarios.

Input Type
Output Type
Input$0.3/M tokens
Output$1.2/M tokens
Context-
Max Output-
6.47Btokens
99.43%Cache Hit Rate

MiniMax-M2.5 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

Compared to its predecessor, M2.1 delivers cleaner, more concise outputs and faster perceived response times. It shows leading multilingual coding performance across major systems and application languages, achieving 49.4% on Multi-SWE-Bench and 72.5% on SWE-Bench Multilingual, and serves as a versatile agent “brain” for IDEs, coding tools, and general-purpose assistance.

To avoid degrading this model's performance, MiniMax highly recommends preserving reasoning between turns.

Input Type
Output Type
Input$0.3/M tokens
Output$1.2/M tokens
Context-
Max Output-
14.53Btokens
93.21%Cache Hit Rate

M2.7 delivers outstanding performance in real-world software engineering, including end-to-end complete project delivery, log analysis and bug triaging, code security, machine learning, and more. On the benchmark SWE-Pro, M2.7 scores 56.22%, nearly matching the level of Opus. This capability also extends to end-to-end complete project delivery scenarios (VIBE-Pro 55.6%) and deep understanding of complex engineering systems on Terminal Bench 2 (57.0%).

In the professional office domain, we have improved the model's specialized knowledge and task delivery capabilities across various fields. On GDPval-AA, its ELO score is 1495, the highest among open-source models. M2.7's ability to perform complex editing in the Office suite (Excel/PPT/Word) has significantly improved, enabling better multi-round revisions and high-fidelity editing. M2.7 is capable of interacting with complex environments. Across 40 complex skills (> 2000 tokens) cases, M2.7 still maintains a 97% skill adherence rate. In OpenClaw usage, M2.7 has shown significant improvement compared to M2.5, scoring close to the latest Sonnet 4.6 in the MMClaw evaluation.

M2.7 possesses excellent identity retention capabilities and emotional intelligence. Beyond productivity use cases, it also opens up space for innovation in interactive entertainment scenarios.

Input Type
Output Type
Input$0.3/M tokens
Output$1.2/M tokens
Context-
Max Output-
1.46Btokens
94.3%Cache Hit Rate

M2.7 highspeed: Same performance, faster, more agile

Input Type
Output Type
Input$0.611/M tokens
Output$2.4439/M tokens
Context-
Max Output-
24.43Btokens
96.31%Cache Hit Rate

MiniMax M3 is the first open-weights flagship that brings coding, agentic reasoning, million-token context, and native multimodality together in one model. It handles autonomous task decomposition, tool use, and multi-step reasoning with ease — and writes code that's meant to ship, not code that just runs. Built on MiniMax's proprietary Sparse Attention architecture, M3 natively supports million-token context windows for long-horizon agents, large codebases, and long-video understanding. Its multimodality isn't bolted on — it's trained in from day one, with text and vision deeply aligned at the semantic level.

Input Type
Output Type
Input$0.3-0.6/M tokens
Output$1.2-2.4/M tokens
Context-
Max Output-
994.51Mtokens
99.25%Cache Hit Rate

M2.5 highspeed: Same performance, faster, more agile

Input Type
Output Type
Input$0.6/M tokens
Output$2.4/M tokens
Context-
Max Output-