Stable LatentMoE (2.8T total, 896 routed experts with 16 active per token; KDA/MLA hybrid attention at 3
reasoning
1 ratio)
Architectural Profile & Capabilities
GDPval-AA v2: 1668 Elo (#3; Fable 5: 1760, Sol: 1748, Opus 4.8: 1600), AA-Briefcase: 1548 Elo (#2; Fable 5: 1583, Sol: 1495), BrowseComp: 91.2% (#1; Sol: 90.4), Terminal-Bench 2.1: 88.3% (#2; Sol: 88.8), DeepSWE: 67.5% (#3; Sol: 73.0, Fable 5: 70.0), FrontierSWE: 81.2% (#2; Fable 5: 86.6), SWE Marathon: 42.0% (#1; Opus 4.8: 40.0), Program Bench: 77.8% (#1; Sol: 77.6), Automation Bench: 30.8% (#1; Sol: 29.7), SpreadsheetBench 2: 34.8% (#1; Fable 5: 34.7), JobBench: 52.9% (#2; Fable 5: 57.4), CharXiv (RQ) w/ tool: 91.3%, Zerobench w/ tool: 41.0%, Kimi Code Bench 2.0 (internal): 72.9%. *Moonshot internal:* Online Exp: 75.5, DECK-Bench: 73.5, Finance-Bench: 62.6 (all #1 vs Opus 4.8 & GPT-5.5). *Still pending:* SWE-bench Verified, MMLU/MMLU-Pro, MATH-500 | Moonshot AI's most capable model (released Jul 16, 2026). 2.8T-parameter multimodal reasoning model — the largest in the Kimi family. 1M-token context window (4× the K2.x family's 256K) with 75% KV cache reduction and up to 6× decoding throughput at 1M context via Kimi Delta Attention (KDA). Agent Swarm scales to 300 sub-agents / 4,000 actions. Strong across coding (Terminal-Bench 88.3% #2, FrontierSWE 81.2% #2, SWE Marathon 42.0% #1), agentic automation (Automation Bench 30.8% #1, SpreadsheetBench 2 34.8% #1), long-context retrieval (BrowseComp 91.2% #1), financial/legal document analysis, patent indexing, and professional slide generation ("OK Computer"). **Best for:** long-horizon agentic coding (SWE Marathon, multi-file refactors), enterprise automation, 1M-context research/synthesis, BrowseComp-class retrieval, financial analysis, and agent-swarm coordination — comprehensive codebase reviews, multi-document financial audits, patent/legal indexing, and large-scale competitive research. **Caveats: reasoning mode permanently locked to "max" — cannot be reduced, inflating cost and latency on every query. Very slow (28 t/s, 4.02s TTFT). $15.00/M output is the most expensive in the Moonshot lineup (3.75× K2.6). Chatbot Arena early evals (codename "Kivine") showed Claude Fable 5 outperforming on robust UI execution and high-difficulty logic. Inherits K2.6's agentic safety lineage (17/20 on Anthropic's record-tampering eval). Chinese jurisdiction: real-name auth required, non-personal data licensed perpetually to Moonshot AI and affiliates. SWE-bench Verified, MMLU-Pro, and MATH-500 remain pending.**
Architecture Type: Dense Transformer
Standard dense autoregressive transformer architecture with full-attention mechanisms.
Recommended Workloads & Primary Use Cases
long-horizon agentic coding (SWE Marathon
multi-file refactors)
enterprise automation
1M-context research/synthesis
BrowseComp-class retrieval
financial analysis
and agent-swarm coordination — comprehensive codebase reviews
multi-document financial audits
patent/legal indexing
inflating cost and latency on every query
ARMES Zero Data Retention
Calls to Kimi K3 are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Moonshot AI.