GDPval-AA v2: 1668 Elo (#3; Fable 5: 1760, Sol: 1748, Opus 4.8: 1600), AA-Briefcase: 1548 Elo (#2; Fable 5: 1583, Sol: 1495), BrowseComp: 91.2% (#1; Sol: 90.4), Terminal-Bench 2.1: 88.3% (#2; Sol: 88.8), DeepSWE: 67.5% (#3; Sol: 73.0, Fable 5: 70.0), FrontierSWE: 81.2% (#2; Fable 5: 86.6), SWE Marathon: 42.0% (#1; Opus 4.8: 40.0), Program Bench: 77.8% (#1; Sol: 77.6), Automation Bench: 30.8% (#1; Sol: 29.7), SpreadsheetBench 2: 34.8% (#1; Fable 5: 34.7), JobBench: 52.9% (#2; Fable 5: 57.4), CharXiv (RQ) w/ tool: 91.3%, Zerobench w/ tool: 41.0%, Kimi Code Bench 2.0 (internal): 72.9%.
Moonshot AI
Multi-agent reasoning, math, deep research.
Laboratory Overview
Moonshot AI focuses on long-context modeling and multi-turn conversational agents with their Kimi family (K3, K2.7). Highly capable in document analysis, code debugging, and iterative project coordination.
Strict Zero Data Retention (ZDR): Requests are processed ephemerally in RAM. No user data, prompts, or completions are stored, indexed, or used for model training.
Moonshot AI Foundation Models (5)
Available under ARMES unified subscription without separate API keys or billing agreements.
(vendor-reported, in-house suites) Kimi Code Bench v2: 62.0 (+21.8% vs K2.6's 50.9; GPT-5.5 69.0, Opus 4.8 67.4), Program Bench: 53.6 (vs K2.6 48.3), MLS Bench Lite: 35.1 (+31.5% vs K2.6 26.7), Kimi Claw 24/7 Bench: 46.9, MCP Atlas: 76.0 (vs K2.6 69.4), MCP Mark Verified: 81.1 (vs K2.6 72.8) | Cost-efficient long-horizon agentic coding — multi-file refactors, hours-long autonomous engineering, CI/CD, large-codebase analysis across Rust/Go/Python/frontend/DevOps.
SWE-Bench Pro: 58.6 (leading), GPQA Diamond: 90.5%, HLE-Full w/ Tools: 54.0, MathVision (w/ Python): 93.2, GDPval-AA: 1520 Elo (vs 1309 for K2.5), Hallucination rate: 39% (vs K2.5's 65%) | Long-horizon agentic coding, autonomous engineering, full-stack workflows, persistent 24/7 background agents, multi-step DevOps.
GPQA Diamond: 87.6%, SWE-bench: 76.8%, HLE-Full w/ Tools: 50.2% (#1), BrowseComp Swarm: 78.4%, TerminalBench: 50.8% | Research with tools, complex multi-step tasks, agentic workflows, strategic reasoning, document analysis, philosophy, deep conceptual work.
Explore Other AI Labs on ARMES
Compare models across the 21+ leading laboratories unified in your workspace.