More Models from MiniMax
Compare other engines available in this laboratory.
SWE-Bench Pro: 59.0% (#3 public leaderboard; beats GPT-5.5 58.6% & Gemini 3.1 Pro 54.2%, trails Opus 4.7/4.8), SWE-Bench Verified: 85.0%, Terminal-Bench 2.1: 66.0%, MCP Atlas: 74.2%, BrowseComp: 83.5 (beats Opus 4.7's 79.3), OSWorld-Verified: 70.06%, SWE-fficiency: 34.8%, KernelBench Hard: 28.8%, SVG-Bench: > Opus 4.7, OmniDocBench: > Gemini 3.1 Pro | Long-horizon agentic coding, full-project delivery, autonomous engineering, multimodal document/chart/video understanding, computer-use (desktop operation), and million-token long-context reasoning at frontier-class quality for ~5–10% of proprietary-model cost.
PinchBench: 86.2% (5th overall, within 1.2pts of Opus 4.6), SWE-Pro: 56.22%, GDPval-AA: 1495 Elo (highest open-source), MLE Bench Lite: 66.6%, Terminal Bench 2: 57.0% | Agentic coding, full-project delivery, complex engineering systems, autonomous debugging, production incidents, office productivity.
SWE-bench: 80.2% — rivals Claude Opus and GPT-5.2 | Coding, debugging, refactoring, full-stack development, front-end components, test coverage, technical implementation.