More Models from xAI
Compare other engines available in this laboratory.
Artificial Analysis Intelligence Index: 54 (#4 of 168), AA Coding Agent Index: 76 (on par with GPT-5.5), τ³-Banking: 33% (#1 of 28 models), GPQA Diamond: 93%, Terminal-Bench 2.1: 82–83.3%, SWE-Bench Pro: 64.7%, DeepSWE 1.0: 62%, SWE Marathon: 29% (beats Opus 4.8's 26%), ~15,954 output tokens per SWE-Bench Pro task (4.2× more token-efficient than Opus 4.8) | xAI's flagship model for software engineering, agentic execution, and technical reasoning (released Jul 8, 2026).
CaseLaw v2: 79.3% (#1, +25 pts over 4.20), CorpFin: #1, τ²-Bench Telecom: 97.7%, GDPval-AA: 1500 ELO, GPQA Diamond: 90.1%, IFBench: 81.0%, SciCode: 47.3%, Coding Index: 41–42.2%, Terminal-Bench Hard: 38.0%, HLE: 35.0%, Artificial Analysis Intelligence Index: 53 | xAI's most capable reasoning model (beta Apr 17, 2026; GA Apr 30, 2026).
SWE-bench Verified: 78.0%, GPQA Diamond: 82.7%, MMLU-Pro: 83.7%, AIME 2025: 90.2%, IFBench: 72.6%, SciCode: 42.0%, Terminal-Bench Hard: 31.0%, τ²-Bench Telecom: 64.4%, HLE: 22.8–30.0%, non-hallucination: ~78% (Thinking mode) | xAI's large-context multi-agent workhorse (beta Feb 17, 2026; GA Mar 10, 2026).