Terminal-Bench 2.1: 87.9%, DeepSWE: 62.7%, CyberGym: 83.3%, AutomationBench: 31.8%, Toolathlon-Verified: 74.1%, HLE w/ tools: 60.0%, GPQA Diamond: 90.1%, Artificial Analysis Intelligence Index: 53.2, AA Coding Index: 68.8, AA Agentic Index: 49.6 | The official GA release of V4 Pro, superseding the April preview (released Aug 13, 2026).
DeepSeek
Versatile reasoning at the lowest cost.
Laboratory Overview
DeepSeek is an open research powerhouse known for pioneering breakthrough Mixture-of-Experts architectures like Multi-Head Latent Attention (MLA) and DualPipe MoE. Their V4 Pro and R1 models offer frontier math and coding at unprecedented cost efficiency.
Strict Zero Data Retention (ZDR): Requests are processed ephemerally in RAM. No user data, prompts, or completions are stored, indexed, or used for model training.
DeepSeek Foundation Models (6)
Available under ARMES unified subscription without separate API keys or billing agreements.
V4 Pro is a high-performance foundation model engineered by DeepSeek, accessible with Zero Data Retention on ARMES.
Terminal-Bench 2.1: 82.7%, CyberGym: 76.7%, Toolathlon-Verified: 70.3%, DSBench-FullStack: 68.7%, DSBench-Hard: 59.6%, DeepSWE: 54.4%, NL2Repo: 54.2%, Agents' Last Exam: 25.2%, AutomationBench Public: 25.1%, Artificial Analysis Intelligence Index: 50 | The official GA release of V4 Flash, superseding the April preview (released Jul 31, 2026).
Strong on long-document and routine coding | Cost-effective high-volume tasks, long-document processing, routine coding, local deployment on consumer hardware.
GPQA Diamond: 82.4%, SWE-bench: 73.1%, AIME: 93.1% | General chat, simple Q&A, translation, notes retrieval, simple scripts, moderate analysis, standard tasks.
Explore Other AI Labs on ARMES
Compare models across the 21+ leading laboratories unified in your workspace.