AA Intelligence Index: 61 (matches GPT-5.6 Sol), GDPval-AA v2: 1753 Elo, CursorBench v3.2: 69.9%, DeepSWE v1.1: 65.9%, FrontierCode v1.1 Extended: 61.3%, APEX-Agents: 57.5%, Terminal-Bench v3.0: 26%, APEX-SWE: 56.4%, AA-Briefcase: 1577 Elo, Harvey LAB: 15.8%, SWE-bench Verified (Vals): 95.60%, GPQA Diamond: 94.9%, HLE: 42.9%, AA Coding Index: 76.8, AA Agentic Index: 58.7, AA Non-Hallucination Rate: 65.7% | xAI's current frontier model (released Aug 12, 2026; knowledge cutoff Feb 1, 2026).

xAI
π integration, real-time sentiment/trend parsing.
Laboratory Overview
Founded by Elon Musk to understand the universe, xAI develops the Grok series. Known for real-time truth-seeking, mathematical precision, and rapid multi-agent reasoning with ultra-large context capabilities.
Strict Zero Data Retention (ZDR): Requests are processed ephemerally in RAM. No user data, prompts, or completions are stored, indexed, or used for model training.
xAI Foundation Models (4)
Available under ARMES unified subscription without separate API keys or billing agreements.
Artificial Analysis Intelligence Index: 54 (#4 of 168), AA Coding Agent Index: 76 (on par with GPT-5.5), ΟΒ³-Banking: 33% (#1 of 28 models), GPQA Diamond: 93%, Terminal-Bench 2.1: 82β83.3%, SWE-Bench Pro: 64.7%, DeepSWE 1.0: 62%, SWE Marathon: 29% (beats Opus 4.8's 26%), ~15,954 output tokens per SWE-Bench Pro task (4.2Γ more token-efficient than Opus 4.8) | xAI's flagship model for software engineering, agentic execution, and technical reasoning (released Jul 8, 2026).
CaseLaw v2: 79.3% (#1, +25 pts over 4.20), CorpFin: #1, ΟΒ²-Bench Telecom: 97.7%, GDPval-AA: 1500 ELO, GPQA Diamond: 90.1%, IFBench: 81.0%, SciCode: 47.3%, Coding Index: 41β42.2%, Terminal-Bench Hard: 38.0%, HLE: 35.0%, Artificial Analysis Intelligence Index: 53 | xAI's most capable reasoning model (beta Apr 17, 2026; GA Apr 30, 2026).
SWE-bench Verified: 78.0%, GPQA Diamond: 82.7%, MMLU-Pro: 83.7%, AIME 2025: 90.2%, IFBench: 72.6%, SciCode: 42.0%, Terminal-Bench Hard: 31.0%, ΟΒ²-Bench Telecom: 64.4%, HLE: 22.8β30.0%, non-hallucination: ~78% (Thinking mode) | xAI's large-context multi-agent workhorse (beta Feb 17, 2026; GA Mar 10, 2026).
Explore Other AI Labs on ARMES
Compare models across the 21+ leading laboratories unified in your workspace.