Claude Opus 5: Everything You Need to Know
On the morning of July 24, 2026, Anthropic released Claude Opus 5. Within the hour, it was live inside ARMES over zero-data-retention inference — processed and forgotten, like everything else on the platform.
This post is not a product announcement. It is a deep dive into what Opus 5 is, what changed from Opus 4.8, where it excels, where it does not, and what it means for anyone doing serious work with AI.
What is Claude Opus 5?
Claude Opus 5 is Anthropic's premium workhorse model — the highest-capability model in their lineup that most developers and professionals are expected to use day-to-day. It sits one tier below the flagship Claude Fable 5 in Anthropic's hierarchy, but approaches Fable-level intelligence on many benchmarks at roughly half the cost.
The current Anthropic model lineup looks like this:
| Model | Position | Intended Use |
|---|---|---|
| Fable 5 | Flagship (Mythos-class) | Multi-hour autonomous agents, hardest reasoning problems |
| Opus 5 | Premium workhorse | Coding, research, enterprise knowledge work, production agents |
| Sonnet 5 | Mid-tier | General assistant, agents, production apps |
| Haiku 4.5 | Fastest & cheapest | High-volume inference, latency-sensitive tasks |
This is a notable positioning shift. In previous generations, Opus was the flagship. Now Anthropic is saying: "Unless you truly need our absolute best model for multi-hour unsupervised agents, use Opus 5." For most users, most of the time, Opus 5 is the ceiling.
The numbers
Here is how Opus 5 compares across public benchmarks:
| Benchmark | Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|---|
| Agentic terminal coding (Frontier-Bench v0.1) | 43.3% | 33.7% | 21.1% | 34.4% |
| Knowledge work (GDPval-AA v2) | 1861 | 1747 | 1593 | 1736 |
| Novel problem-solving (ARC-AGI-3) | 30.2% | — | 1.5% | 7.8% |
| Agentic search (BrowseComp) | 90.8% | 87.4% | 84.3% | 90.4% |
| Multidisciplinary reasoning (HLE — no tools) | 56.3% | 56.5% | 49.8% | — |
| Multidisciplinary reasoning (HLE — with tools) | 64.7% | 63.9% | 57.9% | — |
| Computer use (OSWorld 2.0) | 70.6% | 66.1% | 55.7% | 62.6% |
| Agentic coding (DeepSWE v1.1) | 68.8% | 69.7% | 59.0% | 72.7% |
| Agentic coding (FrontierCode v1.1) | 53.4% | 53.5% | 46.5% | 47.5% |
| Business workflows (AutomationBench) | 26.0% | 17.4% | 17.0% | 18.1% |
| Legal (Legal Agent Benchmark) | 11.7% | 13.3% | 10.4% | 2.5% |
| Health (HealthBench Professional) | 59.8% | 66.0% | 57.4% | 60.5% |
| Biology (BioMysteryBench — hard) | 49.4% | 46.5% | 42.4% | — |
Bold indicates the best score in each row.
Where Opus 5 excels
1. Agentic terminal coding
This is the headline improvement. On Frontier-Bench v0.1, Opus 5 scores 43.3% — more than double Opus 4.8's 21.1%, and nearly 10 points above the next-best model. On CursorBench 3.2 at max effort, it performs within 0.5% of Fable 5's peak score at half the cost per task.
Anthropic's partners confirm the gains are real:
"Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn't just better on our hardest agentic coding tasks, up 22% over Opus 4.7 — it's steadier, with far less variance run to run." — Lovable
The lower variance is significant. A model that occasionally produces brilliant code but frequently produces broken code is not useful for production workflows. Opus 5 is consistent.
2. Novel reasoning (ARC-AGI-3)
The most surprising result:
- Opus 4.8: 1.5%
- GPT-5.6 Sol: 7.8%
- Opus 5: 30.2%
That is a 20x improvement over its predecessor, and nearly 4x the next-best competitor.
ARC-AGI benchmarks are specifically designed to test reasoning on unfamiliar problems rather than pattern matching against training data. A jump of this magnitude suggests Anthropic made meaningful architectural improvements in general reasoning — not just broader training.
3. Knowledge work
GDPval-AA v2: 1861 — the highest published score from any commercial model.
This benchmark measures professional tasks like analysis, writing, synthesis, report generation, planning, and structured reasoning. Opus 5 leads by over 100 points.
4. Computer use
OSWorld 2.0: 70.6% — best in class.
This measures real operating system interaction: clicking, typing, navigating applications, and completing multi-step workflows. Opus 5 outperforms every other model at any given cost, surpassing even Fable 5's best result at roughly a third of the cost.
5. Business automation
AutomationBench: 26.0% — about 1.5x the next-best model.
Even at its lowest effort setting, Opus 5 passes more business automation tasks than any other model at any effort level. This suggests deep optimization for the kind of multi-tool orchestration that enterprise workflows demand.
6. Financial reasoning
Partner testimonials are particularly strong here:
"On some of our hardest financial-modeling tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both accuracy and efficiency. Its performance floor is materially higher, especially on deep finance domain logic. Across effort levels it averaged 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time."
"Claude Opus 5 is the strongest Opus model we've tested on our trading benchmark, and it gets there using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8. Better answers at a fraction of the compute."
Where it does not lead
Transparency matters. Opus 5 is not the best model at everything:
- DeepSWE v1.1: GPT-5.6 Sol leads at 72.7% vs. Opus 5's 68.8%.
- HealthBench Professional: Fable 5 (via Mythos 5) leads at 66.0% vs. 59.8%.
- Legal Agent Benchmark: Fable 5 leads at 13.3% vs. 11.7%.
- HLE without tools: Fable 5 edges Opus 5 by 0.2 points.
If you need absolute maximum capability on multi-hour autonomous agents or health-domain reasoning, Fable 5 remains the answer. But for the vast majority of use cases — coding, research, business automation, knowledge work — Opus 5 matches or exceeds it.
Technical specifications
| Parameter | Value |
|---|---|
| Model ID | claude-opus-5 |
| Context window | 1,000,000 tokens (default and maximum) |
| Max output tokens | 128,000 |
| Training data cutoff | May 2026 |
| Knowledge cutoff | May 2026 |
| Thinking | On by default (adaptive) |
| Effort levels | low, medium, high, xhigh, max |
| Default effort | high |
| Input pricing | $5.00 per million tokens |
| Output pricing | $25.00 per million tokens |
| Prompt caching | Up to 90% cost savings |
| Batch processing | 50% cost savings |
New capabilities
The effort ladder
Opus 5 supports five effort levels, giving you explicit control over how hard the model works:
- Low / Medium — Stronger on Opus 5 than on earlier Opus models. Use as your primary cost and latency controls.
- High — The default. Appropriate for most intelligence-sensitive workloads.
- xHigh — Recommended for coding and agentic work.
- Max — Unconstrained token spending for capability-critical work. Start with 64k max_tokens and tune from there.
A critical behavior change: on Opus 5, thinking cannot be disabled at xhigh or max effort. Requests that set thinking: {"type": "disabled"} at those levels return a 400 error. This is a breaking change from Opus 4.8.
Adaptive thinking
Thinking is enabled by default on Opus 5. The model decides when and how deeply to reason based on the complexity of each turn. At higher effort levels, it thinks on most requests and at greater length. At lower levels, it can skip thinking entirely for simpler problems.
This means you are paying for reasoning tokens only when the model determines they add value — not on every request unconditionally.
Fast mode
A premium pricing option that prioritizes lower latency for applications needing quick responses. Delivers up to 2.5x higher output speed.
Mid-conversation tool changes (beta)
A new capability that lets you add or remove tools between turns of a conversation while preserving the prompt cache. This is significant for agentic workflows where the available tool set evolves as the task progresses.
Efficiency improvements
One of the underreported aspects of Opus 5 is how much more efficient it is:
- On legal work: achieves similar performance to Opus 4.8 while generating 26% fewer tokens on average at max reasoning.
- On trading benchmarks: uses roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8 for the same (or better) answers.
- On financial modeling: achieves 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time.
The pattern is clear: Opus 5 does not just produce better results — it produces them with less waste. This is critical for cost-sensitive production deployments.
Pricing context
Anthropic kept pricing identical to Opus 4.8:
- $5 / million input tokens
- $25 / million output tokens
To put this in perspective:
- Fable 5 costs roughly 2x more.
- GPT-5.6 Sol is comparably priced but Opus 5 outperforms it across most benchmarks.
- With prompt caching (up to 90% savings) and batch processing (50% savings), effective costs can be significantly lower.
For the capability delivered, this is among the strongest price-to-performance ratios in frontier AI.
Availability
Claude Opus 5 is available immediately on:
- Claude API (
claude-opus-5) - Claude in Amazon Bedrock
- Claude on Google Cloud (Vertex AI)
- Claude in Microsoft Foundry
- Claude Pro, Max, Team, and Enterprise plans
And as of this morning: ARMES — with zero-data-retention inference.
Using Opus 5 in ARMES
Opus 5 is available on the Ultra plan. You can:
- Select it manually from the model picker in any chat or Command Center conversation.
- Let Auto routing choose it — the router will direct your most demanding tasks to Opus 5 when the complexity warrants it.
Like every model on ARMES, Opus 5 runs through our zero-data-retention infrastructure. Your prompts are sent to the model, the response is generated, and nothing is stored by the inference provider. No training on your data. No logs retained. The same privacy commitment you trust for every other interaction.
Who should use it
Developers and engineers: If you are writing code with AI — especially autonomous coding agents, complex debugging, or large-scale refactoring — Opus 5 is the best option available at this price point.
Researchers and analysts: The knowledge work and reasoning improvements make it ideal for literature review, data synthesis, financial modeling, and scientific analysis.
Enterprise teams: Business automation scores suggest optimization for the multi-tool orchestration that real workflows demand. Documents, spreadsheets, CRM automation, internal tooling.
Anyone on Opus 4.8: Opus 5 is strictly better. Same price. Better results. Less waste.
The bigger picture
Opus 5 represents a shift in how Anthropic thinks about model tiers. Rather than a simple capability ladder, they are building models optimized for different usage patterns:
- Fable 5 for unsupervised, multi-hour autonomous work where you fire-and-forget.
- Opus 5 for the day-to-day premium workload — the model you open when the task matters and you want the best practical tool.
- Sonnet 5 for scale and speed when you need good-enough at lower cost.
For most people reading this, Opus 5 is the model that matters. It is the one you will reach for when the stakes are high, the problem is hard, and you want something that thinks before it writes.
It is available now. Go use it.
Claude Opus 5 is available on ARMES Ultra. Start free, explore the platform, and upgrade when you are ready for frontier intelligence with zero data retention.
Written by
ARMES Team
From the team building ARMES — private AI that puts every frontier model in one place.