More Models from xAI
Compare other engines available in this laboratory.
AA Intelligence Index: 61 (matches GPT-5.6 Sol), GDPval-AA v2: 1753 Elo, CursorBench v3.2: 69.9%, DeepSWE v1.1: 65.9%, FrontierCode v1.1 Extended: 61.3%, APEX-Agents: 57.5%, Terminal-Bench v3.0: 26%, APEX-SWE: 56.4%, AA-Briefcase: 1577 Elo, Harvey LAB: 15.8%, SWE-bench Verified (Vals): 95.60%, GPQA Diamond: 94.9%, HLE: 42.9%, AA Coding Index: 76.8, AA Agentic Index: 58.7, AA Non-Hallucination Rate: 65.7% | xAI's current frontier model (released Aug 12, 2026; knowledge cutoff Feb 1, 2026).
Artificial Analysis Intelligence Index: 54 (#4 of 168), AA Coding Agent Index: 76 (on par with GPT-5.5), τ³-Banking: 33% (#1 of 28 models), GPQA Diamond: 93%, Terminal-Bench 2.1: 82–83.3%, SWE-Bench Pro: 64.7%, DeepSWE 1.0: 62%, SWE Marathon: 29% (beats Opus 4.8's 26%), ~15,954 output tokens per SWE-Bench Pro task (4.2× more token-efficient than Opus 4.8) | xAI's flagship model for software engineering, agentic execution, and technical reasoning (released Jul 8, 2026).
SWE-bench Verified: 78.0%, GPQA Diamond: 82.7%, MMLU-Pro: 83.7%, AIME 2025: 90.2%, IFBench: 72.6%, SciCode: 42.0%, Terminal-Bench Hard: 31.0%, τ²-Bench Telecom: 64.4%, HLE: 22.8–30.0%, non-hallucination: ~78% (Thinking mode) | xAI's large-context multi-agent workhorse (beta Feb 17, 2026; GA Mar 10, 2026).