More Models from NVIDIA
Compare other engines available in this laboratory.
SWE-Bench Verified: 70.7–71.9%, PinchBench: 90.0%, RULER @1M: 94.7%, LiveCodeBench v6: 89.0%, IOI 2025: 570, GPQA (no tools): 87.0%, IFBench: 81.7%, IMOAnswerBench: 88.6–92.3%, Terminal-Bench 2.1: 56.4%, Artificial Analysis Intelligence Index: 48 (highest US open model) | NVIDIA's most capable model (released Jun 4, 2026).
GPQA Diamond: 79.23% (82.70% w/ tools), HMMT Feb 2025: 93.67% (94.73% w/ tools), AIME 2025: 90.21%, MMLU-Pro: 83.73%, LiveCodeBench v5: 81.19%, SWE-Bench Verified: 60.47%, SWE-Bench Multilingual: 45.78%, RULER @1M: 91.75% (vs GPT-OSS-120B's 22.30), RULER @256K: 96.30%, AA Intelligence Index: 36 | Agentic reasoning, tool calling, long-context workflows, instruction following, multi-agent orchestration.