Smarter Routing: How We Rebuilt Every ARMES Tier This Month
Every query you send to ARMES passes through a routing agent that picks the best model for your task. It looks at what you are asking, considers the strengths and weaknesses of every model in your tier's pool, and sends your prompt to the one most likely to deliver a great answer on the first attempt.
This month we took a hard look at every tier. We questioned every model's inclusion, searched for bias left over from earlier versions, tested new entrants from labs around the world, and rebuilt the routing logic from scratch. The goal was simple: make sure every model in every pool actually earns its slot today, not because it earned it six months ago.
Here is what changed and why.
The philosophy: intelligence-first, route down
Across all four tiers we adopted the same core principle. Start with the most capable model the tier can offer, and only route to a lighter or more specialized model when the task clearly benefits from it. Previous versions of the router defaulted to safe, familiar models and escalated to stronger ones when it detected complexity. That meant users were often getting less intelligence than they could have been.
We flipped it. The default is now the smartest model in the pool. The router routes down to specialists when their specific strengths matter more than raw intelligence, or when the task is simple enough that a lighter model handles it just as well.
Ultra: five models, zero compromises
Pool: Grok 4.6, GPT-5.6 Sol, Claude Opus 4.8, Gemini 3.7 Flash, GPT-5.6 Luna
The Ultra tier went from nine models to five. The previous version spread traffic across too many overlapping options. Models like Sonnet 4.6, Haiku 4.5, GPT-5.5, and Terra were all doing similar jobs at different capability levels, and the router had to make increasingly arbitrary distinctions between them.
Grok 4.6 from xAI is the new default. It scores 61 on the Artificial Analysis Intelligence Index, tying with GPT-5.6 Sol for the highest measured intelligence of any model available through ARMES. It handles engineering, architecture, strategy, debugging, knowledge work, and general reasoning. If you ask something and the router is not sure which specialist to pick, it goes to Grok.
GPT-5.6 Sol from OpenAI is the expanded peak-intelligence flagship. It was underused in the previous version. Sol is now the primary model for the hardest problems in any domain: formal mathematics, deep debugging, advanced science, cybersecurity, complex financial modeling, and multi-agent orchestration.
Claude Opus 4.8 from Anthropic handles creative writing, emotional support, legal and policy reasoning, careful code review, and premium communications. It is also the only model in the pool with computer-use and browser-agent capabilities.
Gemini 3.7 Flash from Google is the fast-intelligence workhorse. Released August 13, it scores 56 on the Intelligence Index with massive coding and agentic gains over its predecessor. It handles research, current information, science, quantitative analysis, rapid coding, multilingual tasks, and large-context ingestion.
GPT-5.6 Luna from OpenAI is the lightweight tier for mechanical and trivial tasks: extraction, classification, formatting, simple greetings, and quick drafts.
Pro: the biggest philosophy shift
Pool: Grok 4.6, Gemini 3.7 Flash, Claude Sonnet 4.6, GPT-5.6 Luna
Pro saw the most dramatic rethinking. The previous router treated Claude Sonnet as the trusted default and routed up to stronger models for harder tasks. That made sense when Sonnet was the clear best model at Pro-tier capability. It no longer is.
Grok 4.6 takes over as the default intelligence anchor for Pro, with the same Intelligence Index score of 61 that makes it the Ultra default. Every engineering, debugging, architecture, analysis, strategy, and knowledge-work query now starts on Grok.
Gemini 3.7 Flash becomes the fast-intelligence workhorse for research, science, quantitative analysis, multilingual work, and rapid coding. Its Intelligence Index of 56 and GPQA Diamond score of 94.5% make it the strongest science and STEM model in the Pro pool.
Claude Sonnet 4.6 keeps a focused lane: creative writing, brand voice, emotional support, legal and policy reasoning, code review requiring judgment, and interpersonal communication. These are genuine Sonnet strengths where no other model in the pool matches it.
GPT-5.6 Luna handles extraction, classification, formatting, batch processing, casual greetings, and quick drafts.
We also cut Haiku 4.5 from the pool entirely. It was a holdover from earlier versions that kept it for its speed and low hallucination rate, but Gemini 3.7 Flash now matches or exceeds it on both calibration and intelligence at comparable throughput.
Eco: energy-efficient intelligence, rebuilt
Pool: DeepSeek V4 Flash 0731, MiniMax M3, Kimi K2.6, GLM-5.2, DeepSeek V4 Pro 0813, Nemotron 3 Ultra, Gemini 3.7 Flash, Qwen 3.8 Max
The Eco tier has a distinct identity: every model activates only a fraction of its total parameters per response. This sparse Mixture-of-Experts approach delivers strong capability while using dramatically less compute and energy per answer. The pool spans a deliberate gradient from ultra-light models (DeepSeek V4 Flash with 13 billion active parameters) through mid-weight specialists up to the heavyweight intelligence ceiling.
Three major additions define this update:
Qwen 3.8 Max from Alibaba is the new peak-intelligence ceiling. It is the highest-scoring open-weight model on the Artificial Analysis Intelligence Index at 58, with 95 billion active parameters out of 2.4 trillion total. PaperBench 93.0. Terminal-Bench 86.6. SWE-bench Pro 67.7. It handles the hardest cross-domain problems, complex multi-file coding, and frontier-level deliverables. The router sends it roughly 5% of volume, only when lighter models are clearly insufficient.
DeepSeek V4 Pro 0813 is the GA release of DeepSeek's V4 Pro, upgrading from the April preview. The Intelligence Index rose from 45 to 53. LiveCodeBench 93.5% and a Codeforces rating of 3206 make it the strongest pure-math and algorithmic-coding model in the pool. Its lane is narrow but unmatched: mathematics, statistics, quantitative science, abstract logic, and competitive programming.
Gemini 3.7 Flash expands from a narrow tool-orchestration role to a broader calibrated-intelligence anchor. Its GPQA Diamond of 94.5% is the highest in the Eco pool, and it now covers advanced STEM reasoning, PhD-level science, finance and legal analysis, and academic writing alongside its tool-orchestration duties.
The pool also retains DeepSeek V4 Flash 0731 as the ultra-cheap workhorse for general coding and transformational work, MiniMax M3 as the calibrated factual anchor (lowest hallucination rate in the pool), Kimi K2.6 from Moonshot AI for SERP research and multilingual translation, GLM-5.2 from Zhipu AI for frontend coding and deep-knowledge conceptual questions, and NVIDIA Nemotron 3 Ultra as the fast generalist and creative/conversational fallback.
Free: leaner, smarter, four models
Pool: DeepSeek V4 Flash 0731, Gemini 3.5 Flash Lite, MiniMax M3, HY3
The Free tier went from five models to four while raising the intelligence floor across the board.
The biggest change is the addition of Tencent HY3, a 295-billion-parameter Mixture-of-Experts model with 21 billion active parameters. It scores 42 on the Intelligence Index, with GPQA Diamond at 89.7%, BrowseComp at 84.2 for agentic search, and MCP-Atlas at 79.1 for tool orchestration. Its calibration philosophy of "answer when grounded, state when evidence is missing" produced an internal hallucination rate of just 5.4%.
HY3 absorbs the roles of two models we cut:
GPT-5 mini scored just 25 on the Intelligence Index, making it the least intelligent model in the previous pool. Despite being positioned as the "reliability anchor," its AA-Omniscience Non-Hallucination Rate of 43.6% was worse than Gemini Flash Lite already in the pool. HY3 outperforms it on general intelligence, STEM reasoning, coding benchmarks, and factual calibration.
Nemotron 3 Super scored 26 on the Intelligence Index. It was the multi-tool orchestration specialist, but HY3 at Intelligence Index 42 is a 16-point upgrade for the same role with validated tool-calling scores.
The rest of the pool is unchanged: DeepSeek V4 Flash 0731 remains the everyday default and scoped-execution workhorse at Intelligence Index 50, Gemini 3.5 Flash Lite from Google remains the fastest model in the pool for retrieval, fact grounding, and translation, and MiniMax M3 remains the multi-file and diagnostic coding specialist with SWE-bench Pro at 59.0% and the lowest hallucination rate in the pool.
What this means for you
You do not need to do anything. The routing updates are live now across all tiers. Every query you send will automatically be matched to the best model in your tier's pool.
If you use Auto mode (the default), you are already benefiting from these changes. Your queries are being handled by models that are, in many cases, significantly more capable than what was serving them last week.
If you manually select models on Pro or Ultra, all of these models are available in the manual picker as well. The full, up-to-date model roster is always visible on the AI Access page.
New labs, new models
One of the most notable aspects of this update is the breadth of labs represented. The ARMES model pool now draws from xAI (Grok 4.6), OpenAI (GPT-5.6 Sol, Luna), Anthropic (Claude Opus 4.8, Sonnet 4.6), Google (Gemini 3.7 Flash, 3.5 Flash Lite), DeepSeek (V4 Flash, V4 Pro), Alibaba (Qwen 3.8 Max), Tencent (HY3), MiniMax (M3), Moonshot AI (Kimi K2.6), Zhipu AI (GLM-5.2), and NVIDIA (Nemotron 3 Ultra).
We are not loyal to any single lab. Every model earns its place on current, independently validated benchmarks and real-world performance. When a newer model from any lab outperforms an incumbent, we make the switch.
How we evaluate
Every model in the pool is assessed on the same criteria:
- Artificial Analysis Intelligence Index for general capability
- GPQA Diamond for graduate-level STEM reasoning
- SWE-bench for real-world software engineering
- Terminal-Bench for agentic terminal automation
- AA-Omniscience for factual calibration and hallucination rates
- MCP-Atlas and BrowseComp for tool orchestration and agentic search
- IFBench for instruction adherence
We weight independently verified benchmarks far more heavily than vendor-reported numbers. When a model's self-reported scores diverge from independent testing, we note the gap and route accordingly.
Privacy: always zero data retention
Every model in every tier runs through the same zero-data-retention infrastructure. Your prompts and responses are never stored, never used for training, and never shared. This applies equally to every lab and every model in the pool. The routing change does not affect your privacy in any way.
What is next
We continuously monitor model performance and will make adjustments as new releases arrive. The routing logic is versioned and auditable, and we publish the philosophy and pool composition for every tier so you always know what is serving your queries.
If you have feedback on the routing or notice a model that seems like a poor fit for a particular kind of task, let us know. The whole point of this system is to get you the best answer on the first try, and your experience is the most important benchmark of all.
Smarter models. Sharper routing. Every tier.
Written by
ARMES Team
From the team building ARMES — private AI that puts every frontier model in one place.