61.4% (beats Claude Opus 5 58.6% & GPT-5.6 Sol 53.8%)
Harvey LAB
reasoning
10.0% (beats 3.7 Flash 8.8% & Opus 5 6.7%)
Vals Finance Agent V2
agentic
61.4% (beats Opus 5 58.6% & Sol 53.8%)
HLE-Verified
reasoning
54.9% (beats Sol 54.5% & Opus 5 54.4%)
BioMysteryBench (Human Difficult)
reasoning
56.5% (beats Opus 5 49.4% & Sol 44.7%)
CharXiv (chart reasoning)
reasoning
86.2% (vs 3.7 Flash 84.5%)
LVBench
reasoning
87.8% (agentic mode)
Cybersecurity (Cyber capabilities)
cyber
CWE-Bench 47.2% pass@1
Architectural Profile & Capabilities
Google's next-generation workhorse model (released Sep 2, 2026; codenamed "Skimaki"). Rapid 3-week iteration following 3.7 Flash, delivering architectural refinements, reduced output verbosity, and drastically lower latency. Preferred by Google engineers internally over Claude Opus on the Jetski coding platform for real-world developer workflows. Matches or exceeds frontier models (Opus 5, GPT-5.6 Sol) across DeepSWE (61.4%), Terminal-Bench 2.1 (90.8%), Vals Finance Agent (61.4%), and HLE-Verified (54.9%) at Flash economics. Configurable thinking levels. Multimodal input (text, image, audio, video, PDF). Specialized cybersecurity patching and vulnerability discovery capabilities. Introductory pricing $0.75/$3.75 through Dec 31, 2026 ($1.50/$7.50 thereafter); 90% cache discount ($0.075/M cached). Best for: agentic coding, automated debugging and patching, complex enterprise workflows, finance/legal analysis, scientific/biomedical research, and high-frequency tool use loops
Architecture Type: Mixture of Experts (MoE)
Sparse Mixture-of-Experts routing: activates only a subset of parameter experts per token, delivering frontier-level intelligence with exceptional efficiency.
Recommended Workloads & Primary Use Cases
Everyday reasoning
Drafting
Fast problem solving
ARMES Zero Data Retention
Calls to Gemini 3.8 Flash are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Google.