More Models from Moonshot AI
Compare other engines available in this laboratory.
GDPval-AA v2: 1668 Elo (#3; Fable 5: 1760, Sol: 1748, Opus 4.8: 1600), AA-Briefcase: 1548 Elo (#2; Fable 5: 1583, Sol: 1495), BrowseComp: 91.2% (#1; Sol: 90.4), Terminal-Bench 2.1: 88.3% (#2; Sol: 88.8), DeepSWE: 67.5% (#3; Sol: 73.0, Fable 5: 70.0), FrontierSWE: 81.2% (#2; Fable 5: 86.6), SWE Marathon: 42.0% (#1; Opus 4.8: 40.0), Program Bench: 77.8% (#1; Sol: 77.6), Automation Bench: 30.8% (#1; Sol: 29.7), SpreadsheetBench 2: 34.8% (#1; Fable 5: 34.7), JobBench: 52.9% (#2; Fable 5: 57.4), CharXiv (RQ) w/ tool: 91.3%, Zerobench w/ tool: 41.0%, Kimi Code Bench 2.0 (internal): 72.9%.
SWE-Bench Pro: 58.6 (leading), GPQA Diamond: 90.5%, HLE-Full w/ Tools: 54.0, MathVision (w/ Python): 93.2, GDPval-AA: 1520 Elo (vs 1309 for K2.5), Hallucination rate: 39% (vs K2.5's 65%) | Long-horizon agentic coding, autonomous engineering, full-stack workflows, persistent 24/7 background agents, multi-step DevOps.
GPQA Diamond: 87.6%, SWE-bench: 76.8%, HLE-Full w/ Tools: 50.2% (#1), BrowseComp Swarm: 78.4%, TerminalBench: 50.8% | Research with tools, complex multi-step tasks, agentic workflows, strategic reasoning, document analysis, philosophy, deep conceptual work.