More Models from Moonshot AI
Compare other engines available in this laboratory.
GDPval-AA v2: 1668 Elo (#3; Fable 5: 1760, Sol: 1748, Opus 4.8: 1600), AA-Briefcase: 1548 Elo (#2; Fable 5: 1583, Sol: 1495), BrowseComp: 91.2% (#1; Sol: 90.4), Terminal-Bench 2.1: 88.3% (#2; Sol: 88.8), DeepSWE: 67.5% (#3; Sol: 73.0, Fable 5: 70.0), FrontierSWE: 81.2% (#2; Fable 5: 86.6), SWE Marathon: 42.0% (#1; Opus 4.8: 40.0), Program Bench: 77.8% (#1; Sol: 77.6), Automation Bench: 30.8% (#1; Sol: 29.7), SpreadsheetBench 2: 34.8% (#1; Fable 5: 34.7), JobBench: 52.9% (#2; Fable 5: 57.4), CharXiv (RQ) w/ tool: 91.3%, Zerobench w/ tool: 41.0%, Kimi Code Bench 2.0 (internal): 72.9%.
(vendor-reported, in-house suites) Kimi Code Bench v2: 62.0 (+21.8% vs K2.6's 50.9; GPT-5.5 69.0, Opus 4.8 67.4), Program Bench: 53.6 (vs K2.6 48.3), MLS Bench Lite: 35.1 (+31.5% vs K2.6 26.7), Kimi Claw 24/7 Bench: 46.9, MCP Atlas: 76.0 (vs K2.6 69.4), MCP Mark Verified: 81.1 (vs K2.6 72.8) | Cost-efficient long-horizon agentic coding — multi-file refactors, hours-long autonomous engineering, CI/CD, large-codebase analysis across Rust/Go/Python/frontend/DevOps.
GPQA Diamond: 87.6%, SWE-bench: 76.8%, HLE-Full w/ Tools: 50.2% (#1), BrowseComp Swarm: 78.4%, TerminalBench: 50.8% | Research with tools, complex multi-step tasks, agentic workflows, strategic reasoning, document analysis, philosophy, deep conceptual work.