More Models from NVIDIA
Compare other engines available in this laboratory.
GPQA Diamond: 79.23% (82.70% w/ tools), HMMT Feb 2025: 93.67% (94.73% w/ tools), AIME 2025: 90.21%, MMLU-Pro: 83.73%, LiveCodeBench v5: 81.19%, SWE-Bench Verified: 60.47%, SWE-Bench Multilingual: 45.78%, RULER @1M: 91.75% (vs GPT-OSS-120B's 22.30), RULER @256K: 96.30%, AA Intelligence Index: 36 | Agentic reasoning, tool calling, long-context workflows, instruction following, multi-agent orchestration.
SWE-Bench Verified: 51.56% (BF16) / 52.80% (NVFP4), PinchBench: 85.37%, GPQA Diamond: 75.44%, MMLU Pro: 81.94%, Terminal-Bench 2.1: 24.58%, HLE: 11.72%, IFBench (loose): 71.88%, BrowseComp: 36.97%, AA Intelligence Index: 23.6, AA Coding Index: 26.8, AA Agentic Index: 13.8, AA Non-Hallucination Rate: 62.4% | NVIDIA's highest-efficiency model, purpose-built for the execution layer of always-on agents (released Aug 11, 2026).