Sonnet 5.5
Claude Sonnet 5.5 is Anthropic's Sonnet-class workhorse for well-scoped everyday software work — a 1M-context, image-capable, OpenAI-compatible model priced at $2/$10 per million tokens. It is the mid-tier member of the Claude 5.5 family announced alongside Opus 5.5 (September 22, 2026), designed as the cheap-but-capable successor to Claude Sonnet 5 available over Zero Data Retention inference on ARMES AI.
Empirical Evaluation Results
Architectural Profile & Capabilities
Claude Sonnet 5.5 is Anthropic's Sonnet-tier model positioned between the flagship Opus 5.5 and the light Haiku 5.5 in the Claude 5.5 generation. Anthropic announced the Claude 5.5 family on September 22, 2026 when it launched Claude Opus 5.5, stating that "Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks." The OpenRouter listing (1M-token context, image modality, $2/$10 per-million-token pricing identical to Claude Sonnet 5) indicates the model is a direct upgrade of the Sonnet 5 line targeted at well-scoped everyday engineering: building features, fixing bugs, and producing production code. The model succeeds Claude Sonnet 5, which Anthropic released on June 30, 2026 with a 1M-token context window, 128K maximum output, adaptive extended-thinking, prompt caching, batch processing, and a January 2026 knowledge cutoff. Sonnet 5.5 retains this API surface (1M context, adaptive thinking, tool use, vision) while presumably refreshing the underlying weights and post-training for higher agentic and coding efficiency. Architecture: Anthropic does not publish parameter counts, expert counts, or attention internals for any Claude model, and no technical report for Sonnet 5.5 discloses them. Based on the Sonnet lineage and Anthropic's disclosed behaviors — autoregressive token-by-token generation, extended/adaptive thinking with variable reasoning-token budgets, and a 1M-token context window implying a hybrid local/global attention scheme (sliding-window plus sparse global layers, likely combined with grouped-query attention for KV-cache efficiency) — it is best characterized as a dense autoregressive Transformer rather than a publicly confirmed Mixture-of-Experts. No evidence supports an MoE or diffusion (dLLM) architecture; Anthropic has not released MoE models under the Claude brand. Benchmarks: Sonnet 5.5-specific official scorecards had not been published at the time of research; the model was pre-launch/imminent. The closest verified anchors are its predecessor (Claude Sonnet 5), which third-party trackers report at ~92.4% on SWE-bench Verified and 82.4% on LiveCodeBench, and Anthropic's Terminal-Bench 4.0 chart for the Opus 5.5 launch cohort (Opus 5.5 66.4%, Fable 5.1 55.8%, Opus 5 52.3%). One uncorroborated third-party report cites a Sonnet 5.5 Terminal-Bench 4.0 score of ~70.6%. These figures should be treated as directional, not vendor-verified. Speed: the Sonnet tier is Anthropic's fast, high-volume workhorse class (Anthropic qualitatively designates it "fast"). With adaptive thinking enabled, effective throughput drops as hidden reasoning tokens consume the budget. Best estimate: high-speed output (~90–130 tps non-reasoning), with reasoning-heavy turns effectively landing in the 45–70 tps range and elevated TTFT on very long (multi-hundred-K) contexts. Status note: because Sonnet 5.5 is a just-announced/imminent Claude 5.5-family model, several specifications (exact benchmark scores, parameterization, context-degradation curves) remain officially undisclosed. Verified Sonnet 5 figures are carried forward as the strongest available proxy.
Sonnet-class Anthropic model, successor to Claude Sonnet 5 (released June 30, 2026). Autoregressive decoder with a 1M-token context window and 128K maximum output. Extended/adaptive "thinking" enables variable test-time reasoning-token budgets that are user/API-controllable rather than fixed. A 1M context strongly implies hybrid attention (sliding-window local attention interleaved with sparse global layers) plus grouped-query attention (GQA) for KV-cache efficiency, though Anthropic does not publish MLA/GQA/sliding-window specifics or layer/head/parameter counts. No evidence of Mixture-of-Experts or diffusion (dLLM) — this is not a sparse MoE model. Multimodal (image input) and native tool-use/agentic function calling. Anthropic has not released a technical report or arXiv paper for Sonnet 5.5.
Recommended Workloads & Primary Use Cases
- •Pre-launch/imminent status: Sonnet 5.5 was announced as "following in the coming weeks" after Opus 5.5 (September 22, 2026); official Sonnet 5.5 model card and vendor-published benchmark scorecard were not yet available at research time, so scores cited are proxies from Claude Sonnet 5 and third-party trackers.
- •No published architecture disclosure: Anthropic does not publish parameter counts, MoE/dense confirmation, attention type (MLA/GQA/sliding-window), or exact reasoning-token budgets, so capacity planning must rely on empirical latency/cost testing rather than spec sheets.
- •Adaptive-thinking latency tax: enabling extended thinking materially raises time-to-first-token and lowers effective tokens/second; hidden reasoning tokens are billed as output, so cost per task can exceed naive per-token estimates by 2-5x on hard reasoning.
- •Long-context degradation: near the 1M-token ceiling, expect "lost-in-the-middle" recall loss and instruction-adherence drift; keep critical instructions at the head/tail and validate retrieval accuracy above ~200K tokens. Max output caps at 128K tokens.
- •Prompt-formatting sensitivity: Claude models favor explicit XML-style structure tags and clear role/tool delimiting; ambiguous or malformed tool schemas increase refusal/retry loops. Cached-input pricing ($0.20/1M) requires prompt-prefix stability, so reorder dynamic content to the tail.
- •Ecosystem/version skew: benchmark figures are not cross-comparable across harnesses or versions (e.g., Terminal-Bench 2.1 vs 4.0), and vendor system prompts, effort settings, and retry policies shift results — re-benchmark on your own task distribution.
Calls to Sonnet 5.5 are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Anthropic.
More Models from Anthropic
Compare other engines available in this laboratory.
Claude Opus 5.5 is Anthropic's September 2026 flagship Opus-class model for long-horizon agentic coding, computer use and knowledge work — 1M-token context, native image input, adaptive (non-disableable) reasoning at four effort tiers, and $4/$20 per 1M tokens (40% cheaper per typical task than Opus 5). Available with Zero Data Retention inference on ARMES AI.
Anthropic's flagship 1M-context frontier multimodal foundation model, engineered for deep reasoning, autonomous coding agents, and complex long-horizon system execution.
Anthropic's most capable model (released May 28, 2026; knowledge cutoff Jan 2026).