More Models from xAI
Compare other engines available in this laboratory.
AA Intelligence Index: 61 (matches GPT-5.6 Sol), GDPval-AA v2: 1753 Elo, CursorBench v3.2: 69.9%, DeepSWE v1.1: 65.9%, FrontierCode v1.1 Extended: 61.3%, APEX-Agents: 57.5%, Terminal-Bench v3.0: 26%, APEX-SWE: 56.4%, AA-Briefcase: 1577 Elo, Harvey LAB: 15.8%, SWE-bench Verified (Vals): 95.60%, GPQA Diamond: 94.9%, HLE: 42.9%, AA Coding Index: 76.8, AA Agentic Index: 58.7, AA Non-Hallucination Rate: 65.7% | xAI's current frontier model (released Aug 12, 2026; knowledge cutoff Feb 1, 2026).
Artificial Analysis Intelligence Index: 54 (#4 of 168), AA Coding Agent Index: 76 (on par with GPT-5.5), τ³-Banking: 33% (#1 of 28 models), GPQA Diamond: 93%, Terminal-Bench 2.1: 82–83.3%, SWE-Bench Pro: 64.7%, DeepSWE 1.0: 62%, SWE Marathon: 29% (beats Opus 4.8's 26%), ~15,954 output tokens per SWE-Bench Pro task (4.2× more token-efficient than Opus 4.8) | xAI's flagship model for software engineering, agentic execution, and technical reasoning (released Jul 8, 2026).
CaseLaw v2: 79.3% (#1, +25 pts over 4.20), CorpFin: #1, τ²-Bench Telecom: 97.7%, GDPval-AA: 1500 ELO, GPQA Diamond: 90.1%, IFBench: 81.0%, SciCode: 47.3%, Coding Index: 41–42.2%, Terminal-Bench Hard: 38.0%, HLE: 35.0%, Artificial Analysis Intelligence Index: 53 | xAI's most capable reasoning model (beta Apr 17, 2026; GA Apr 30, 2026).