SWE-Bench Verified: 70.7–71.9%, PinchBench: 90.0%, RULER @1M: 94.7%, LiveCodeBench v6: 89.0%, IOI 2025: 570, GPQA (no tools): 87.0%, IFBench: 81.7%, IMOAnswerBench: 88.6–92.3%, Terminal-Bench 2.1: 56.4%, Artificial Analysis Intelligence Index: 48 (highest US open model) | NVIDIA's most capable model (released Jun 4, 2026).
NVIDIA
Agentic speed, hybrid architecture, 1M context.
Laboratory Overview
NVIDIA builds hardware-optimized transformer and hybrid architectures like Nemotron 3 Ultra and Lightning, utilizing Mamba-Transformer hybrid state-spaces for hyper-efficient inference.
Strict Zero Data Retention (ZDR): Requests are processed ephemerally in RAM. No user data, prompts, or completions are stored, indexed, or used for model training.
NVIDIA Foundation Models (3)
Available under ARMES unified subscription without separate API keys or billing agreements.
GPQA Diamond: 79.23% (82.70% w/ tools), HMMT Feb 2025: 93.67% (94.73% w/ tools), AIME 2025: 90.21%, MMLU-Pro: 83.73%, LiveCodeBench v5: 81.19%, SWE-Bench Verified: 60.47%, SWE-Bench Multilingual: 45.78%, RULER @1M: 91.75% (vs GPT-OSS-120B's 22.30), RULER @256K: 96.30%, AA Intelligence Index: 36 | Agentic reasoning, tool calling, long-context workflows, instruction following, multi-agent orchestration.
SWE-Bench Verified: 51.56% (BF16) / 52.80% (NVFP4), PinchBench: 85.37%, GPQA Diamond: 75.44%, MMLU Pro: 81.94%, Terminal-Bench 2.1: 24.58%, HLE: 11.72%, IFBench (loose): 71.88%, BrowseComp: 36.97%, AA Intelligence Index: 23.6, AA Coding Index: 26.8, AA Agentic Index: 13.8, AA Non-Hallucination Rate: 62.4% | NVIDIA's highest-efficiency model, purpose-built for the execution layer of always-on agents (released Aug 11, 2026).
Explore Other AI Labs on ARMES
Compare models across the 21+ leading laboratories unified in your workspace.