Moderate (~46 t/s on Tencent Cloud; native 10B/0.7B MTP layer for speculative decoding)
Architectural Profile & Capabilities
SWE-bench Pro: 65.7%, GPQA Diamond: 92.3%, Terminal-Bench 2.1: 85.4%, DeepSWE v1.1: 64.3%, HLE: 55.4%, SWE-bench Verified: ~82.9%, NL2Repo: 58.9%. Tencent internal blind eval (163 experts, 203 engineering tasks): 2.99/4.00 vs GLM-5.3 2.92, Kimi K3 2.94. BenchLM composite: 79.16/100, estimated public rank #7 | Tencent's next-generation flagship (released Aug 28, 2026). Major leap over Hy3 in scale (295B/21B → 770B/49B) and architecture (traditional attention → Gated DSA + IndexCache). Designed for long-context productivity: coding agents, complex tool-use workflows, office analytics, game development, and scientific research. Contributed to its own development process (recursive self-improvement on training methods, data strategies, and inference optimization -- increased end-to-end throughput 31.8%). Configurable reasoning: `high` (default) and `no_think`. Apache 2.0 licensed, open weights on HuggingFace (BF16 + FP8). **Caveats: preview release with no independent verification (no AA Intelligence Index, no Toolathlon, no MCP-Atlas, no hallucination metrics). Tencent warns of over-reasoning and over-verification in complex tasks (increased latency and token consumption). Day-zero -- monitor for independent evaluation before routing.**
Architecture Type: Mixture of Experts (MoE)
Sparse Mixture-of-Experts routing: activates only a subset of parameter experts per token, delivering frontier-level intelligence with exceptional efficiency.
Recommended Workloads & Primary Use Cases
Everyday reasoning
Drafting
Fast problem solving
ARMES Zero Data Retention
Calls to Hy4 are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Tencent.