Smarter Routing and Expanded Free Access: The September 2026 Platform Upgrade
Every time you type a prompt into ARMES, you are backed by an automated routing intelligence. Rather than forcing you to decide whether your task needs a code specialist, a creative writer, a deep researcher, or a mathematical proof engine, ARMES analyzes your query and pairs it with the single best model in your plan.
Over the past month, we completed an end-to-end audit of our routing architecture across all four tiers: Ultra, Pro, Eco, and Free. We eliminated legacy models that were grandfathered in from earlier seasons, added brand-new frontier models from Google, Alibaba, and Z.ai, and restructured our decision trees so every query receives specialized, near-frontier capability.
Alongside these architectural upgrades, we have also expanded our Free tier experience. Free accounts now receive more than double their previous daily query allowance, supported by an ultra-stable, zero-hallucination model pool.
Here is an overview of what changed, why we changed it, and what it means for your daily workflows.
1. Ultra Tier: The Uncompromising 4-Lab Frontier Quad
Pool: x-ai/grok-4.6, openai/gpt-6-astra, anthropic/claude-opus-4.8, google/gemini-3.8-flash
Our Ultra tier is designed for professionals and teams who require the highest level of intelligence without compromises. In reviewing our previous architecture, we found that routing lighter utility tasks to smaller models diluted the Ultra experience. When you subscribe to an elite tier, you expect world-class intelligence on every single turn.
We consolidated Ultra down to a focused 4-Lab Frontier Quad where every model is an absolute category leader, now anchored by OpenAI's apex generational flagship:
- xAI Grok 4.6 (Default Intelligence Anchor): Scoring 61 on the Artificial Analysis Intelligence Index and holding the top position on GDPval-AA v2 (1753 Elo), Grok 4.6 serves as the primary default. It anchors interactive full-stack implementation, complex architecture design, business strategy, and broad professional reasoning.
- OpenAI GPT-6 Astra (Apex Autonomous Reasoning & Formal Mathematics): OpenAI's latest generational flagship replaces GPT-5.6 Sol entirely on Ultra. Astra delivers an unprecedented leap in capability: 97.6% on FrontierMath Tier 4 (saturating frontier mathematical proofs), 88.0% single-attempt binary reverse engineering on SRE-Bench (vs Sol's 55.9%), 95.9% on BenchCAD geometric code generation, and 57.9% on Terminal-Bench 4.0. Powered by recurrent depth latent reasoning and Codex harness state persistence, Astra enables deep multi-repo distributed systems refactoring while guaranteeing 0.00% sandbox evasion on ExploitGym (down from Sol's 48.2% failure rate).
- Anthropic Claude Opus 4.8 (Writing, Nuance & Human Judgment): The benchmark leader for nuanced prose, creative narratives, executive communications, emotionally sensitive counseling, careful code auditing, and sensitive interpersonal strategy.
- Google Gemini 3.8 Flash (High-Speed Execution, Scoped Code Repair & Science): Newly released on September 2, Gemini 3.8 Flash delivers 305 tokens per second with a 1-million-token window and sub-second time-to-first-token. It handles rapid terminal workflows (Terminal-Bench 90.8%), scoped bug fixes (DeepSWE v1.1 73.8%), automated vulnerability remediation (CWE-Bench 47.2%), quantitative finance (Vals Finance 61.4%), and biomedical research (BioMysteryBench 56.5%). All high-volume data extractions, shell automation, and transformations run through Gemini 3.8 Flash, completely eliminating the need for lesser models on Ultra.
2. Pro Tier: The Curated 6-Powerhouse Pool
Pool: x-ai/grok-4.6, google/gemini-3.8-flash, anthropic/claude-sonnet-4.6, qwen/qwen3.8-2.4t-a95b, deepseek/deepseek-v4-pro-0813, z-ai/glm-5.3-flash
Pro is our most popular tier. To ensure Pro users never experience degraded responses on utility tasks, we eliminated legacy lightweight models and rebuilt the pool around six distinct champions, newly fine-tuned for latency and execution speed:
- Full-Stack Architecture & Default Reasoning:
x-ai/grok-4.6(AA Index 61, CursorBench 69.9% #1, GDPval 1753) anchors end-to-end full-stack features, complex backend services, database schemas, and commercial business strategy. - Rapid Terminal Automation & Scoped Code Repair:
google/gemini-3.8-flash(~305 t/s, sub-second TTFT) takes over all command-line execution, shell scripts, CI/CD pipeline configs, test scaffolding, and rapid bug fixes (DeepSWE v1.1 73.8% — parity with frontier flagships). By routing CLI and repair tasks to Gemini 3.8 Flash, we eliminate the 30–40s initial response latency and avoid terminal bottlenecks on everyday developer workflows. - Creative Prose, Empathy & Legal Care:
anthropic/claude-sonnet-4.6provides best-in-class human warmth, brand voice, persuasive messaging, sensitive counseling, and careful contract review. - Academic Research & Dense Specifications:
qwen/qwen3.8-2.4t-a95bfrom Alibaba leads the world on academic synthesis (PaperBench 93.0 #1) and multi-constraint document authoring (IFBench 82.8). It handles formal whitepapers, RFC-style technical specs, and multi-paper literature reviews across a full 1M context. - Pure Algorithms & Discrete Mathematics:
deepseek/deepseek-v4-pro-0813serves as Pro's pinnacle mathematical anchor (Codeforces 3206 International Grandmaster, LiveCodeBench 93.5% #1 open-weights, GPQA Diamond 90.1%), providing subscribers with near-flagship mathematical proofs, calculus derivations, dynamic programming algorithms, and SQL query-plan optimizations without Ultra-tier costs. - Visual Frontend & High-Speed Extraction:
z-ai/glm-5.3-flash(Code Arena Frontend #2 worldwide) generates modern React, Tailwind, and dashboard interfaces while handling rapid tabular transformations with sub-second responsiveness at $0.50/M.
3. Eco Tier: 100% Native 1M-Context Energy Efficiency
Pool: deepseek/deepseek-v4-flash-0731, z-ai/glm-5.3-flash, minimax/minimax-m3, deepseek/deepseek-v4-pro-0813, nvidia/nemotron-3-ultra-550b-a55b, google/gemini-3.8-flash, qwen/qwen3.8-2.4t-a95b
The Eco tier is built entirely on energy-efficient sparse Mixture-of-Experts architectures. Each model activates only a small fraction of its total parameters per token, delivering remarkable intelligence-per-watt.
In this update, we made two critical structural refinements:
- Removed the 256K Context Bottleneck: We retired earlier models that were capped at 256K context. Every single model in the Eco pool now supports a native 1-million-token context window.
- Elevated Research and Cyber Remediation: Google Gemini 3.8 Flash takes over live web research, active vulnerability remediation, and diagnostic terminal troubleshooting. NVIDIA Nemotron 3 Ultra anchors creative writing and literary translation, while GLM-5.3-Flash leads frontend UI development and deep conceptual research.
4. Free Tier: Tool-Hardened Reliability
Pool: deepseek/deepseek-v4-flash-0731, tencent/hy3, z-ai/glm-5.3-flash
For the Free tier, our goal was simple: ensure that anyone trying ARMES for the first time gets an accurate, impressive, and fast experience without hitting frustrating tool errors or hallucinations.
ARMES agents make frequent use of search and context tools. Some lightweight models struggle with tool calling, outputting raw code syntax or hallucinating arguments. The new Free pool resolves this by deploying three rock-solid workhorses:
- DeepSeek V4 Flash 0731 (Conversational Workhorse): Handles everyday conversation, general writing, text summarization, and single-file utility scripts.
- Tencent HY3 (Grounded Research & Retrieval): Specifically trained to state when evidence is missing rather than fabricate facts, HY3 achieves an industry-leading 5.4 percent hallucination rate, along with 79.1 percent on MCP-Atlas and 84.2 percent on BrowseComp. It handles all live web search queries, current events, and notes retrieval with pinpoint precision.
- Z.ai GLM-5.3-Flash (Peak Intelligence Anchor): Scoring 57 on the Artificial Analysis Intelligence Index, GLM-5.3-Flash brings near-frontier power to our Free tier. It owns multi-file software engineering (DeepSWE 63.4%), terminal debugging (Terminal-Bench 84.3%), and complex STEM problem solving.
5. Expanded Free Daily Limits: 25 Messages per Day
We listened closely to community feedback regarding our free daily usage caps. Previously, users were capped at 10 messages per day with an additional weekly limit. For someone testing an agent on a coding project, drafting a detailed outline, or refining a research query, 10 messages went by too quickly and interrupted productive flow.
Effective immediately:
- Daily allowance increased to 25 messages per day: A 2.5x increase in daily message volume.
- Weekly limit removed: We eliminated the weekly cap entirely. Your allowance refreshes cleanly every 24 hours.
This allows free users to complete two to three full, multi-turn problem-solving sessions every single day. Whether you are debugging code, exploring research with live web tools, or organizing your personal notes, you can now experience the full depth of ARMES without rationing your questions.
Zero Data Retention Across Every Tier
Regardless of whether you are using ARMES on our Free plan, Eco, Pro, or Ultra, our privacy architecture remains identical: Zero Data Retention (ZDR).
Every model in our network processes your requests in memory and immediately forgets them. Your conversations, proprietary code, personal documents, and agent configurations are never stored by model providers and never used to train machine learning models.
These updates are live right now in your workspace. Try out the new models and let us know what you think.
Written by
ARMES Team
From the team building ARMES — private AI that puts every frontier model in one place.