More Models from xAI
Compare other engines available in this laboratory.
AA Intelligence Index: 61 (matches GPT-5.6 Sol), GDPval-AA v2: 1753 Elo, CursorBench v3.2: 69.9%, DeepSWE v1.1: 65.9%, FrontierCode v1.1 Extended: 61.3%, APEX-Agents: 57.5%, Terminal-Bench v3.0: 26%, APEX-SWE: 56.4%, AA-Briefcase: 1577 Elo, Harvey LAB: 15.8%, SWE-bench Verified (Vals): 95.60%, GPQA Diamond: 94.9%, HLE: 42.9%, AA Coding Index: 76.8, AA Agentic Index: 58.7, AA Non-Hallucination Rate: 65.7% | xAI's current frontier model (released Aug 12, 2026; knowledge cutoff Feb 1, 2026).
CaseLaw v2: 79.3% (#1, +25 pts over 4.20), CorpFin: #1, τ²-Bench Telecom: 97.7%, GDPval-AA: 1500 ELO, GPQA Diamond: 90.1%, IFBench: 81.0%, SciCode: 47.3%, Coding Index: 41–42.2%, Terminal-Bench Hard: 38.0%, HLE: 35.0%, Artificial Analysis Intelligence Index: 53 | xAI's most capable reasoning model (beta Apr 17, 2026; GA Apr 30, 2026).
SWE-bench Verified: 78.0%, GPQA Diamond: 82.7%, MMLU-Pro: 83.7%, AIME 2025: 90.2%, IFBench: 72.6%, SciCode: 42.0%, Terminal-Bench Hard: 31.0%, τ²-Bench Telecom: 64.4%, HLE: 22.8–30.0%, non-hallucination: ~78% (Thinking mode) | xAI's large-context multi-agent workhorse (beta Feb 17, 2026; GA Mar 10, 2026).