Terminal-Bench 2.1: 82.7%, CyberGym: 76.7%, Toolathlon-Verified: 70.3%, DSBench-FullStack: 68.7%, DSBench-Hard: 59.6%, DeepSWE: 54.4%, NL2Repo: 54.2%, Agents' Last Exam: 25.2%, AutomationBench Public: 25.1%, Artificial Analysis Intelligence Index: 50 | The official GA release of V4 Flash, superseding the April preview (released Jul 31, 2026). **Same architecture and parameter count — every gain is from re-post-training, not a new model.** Outperforms V4 Pro Preview on 7 of 9 published agent benchmarks despite activating under a third of Pro's parameters (13B vs 49B). Thinking on by default with three reasoning modes (non-think, high, max). Natively supports Responses API with Codex-style agent harness adaptation. DSpark speculative decoding module attached for faster inference. 2,500 concurrent request support (5× V4 Pro). MIT licensed open weights on Hugging Face. **Best for:** high-throughput agentic coding, terminal/shell agent loops, multi-step tool orchestration, security research, data science full-stack tasks, cost-sensitive automation pipelines, and any workload where the ~90× output-price advantage over frontier models matters more than the 2–15 pt capability gap. **Caveats: all benchmark scores are DeepSeek's own internal runs using DeepSeek Harness (unreleased) at max effort — no independent third-party verification at time of writing. NIST CAISI previously found V4 family real-world capability lagging DeepSeek's self-reported figures by ~8 months. Still behind Opus 4.8 on every published benchmark. No system card or safety framework published. Peak/off-peak pricing (2× during Beijing business hours) took effect Aug 16, 2026 — output rose from $0.28 to $1.32 peak / $0.66 off-peak.**
Architecture Type: Mixture of Experts (MoE)
Sparse Mixture-of-Experts routing: activates only a subset of parameter experts per token, delivering frontier-level intelligence with exceptional efficiency.
Recommended Workloads & Primary Use Cases
high-throughput agentic coding
terminal/shell agent loops
multi-step tool orchestration
security research
data science full-stack tasks
cost-sensitive automation pipelines
ARMES Zero Data Retention
Calls to V4 Flash-0731 are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by DeepSeek.