M2.7 delivers outstanding performance in real-world software engineering, including end-to-end complete project delivery, log analysis and bug triaging, code security, machine learning, and more. On the benchmark SWE-Pro, M2.7 scores 56.22%, nearly matching the level of Opus. This capability also extends to end-to-end complete project delivery scenarios (VIBE-Pro 55.6%) and deep understanding of complex engineering systems on Terminal Bench 2 (57.0%).
In the professional office domain, we have improved the model's specialized knowledge and task delivery capabilities across various fields. On GDPval-AA, its ELO score is 1495, the highest among open-source models. M2.7's ability to perform complex editing in the Office suite (Excel/PPT/Word) has significantly improved, enabling better multi-round revisions and high-fidelity editing. M2.7 is capable of interacting with complex environments. Across 40 complex skills (> 2000 tokens) cases, M2.7 still maintains a 97% skill adherence rate. In OpenClaw usage, M2.7 has shown significant improvement compared to M2.5, scoring close to the latest Sonnet 4.6 in the MMClaw evaluation.
M2.7 possesses excellent identity retention capabilities and emotional intelligence. Beyond productivity use cases, it also opens up space for innovation in interactive entertainment scenarios.