Haiku 5.5
Claude Haiku 5.5 is Anthropic's small-model tier for high-volume, latency-sensitive and agentic workloads, shipped Oct 7, 2026 with a 1M-token context window, 128K max output, hybrid reasoning, and aggressive $0.10/$0.50 per-Mtok pricing (sub-100K prompts).
Empirical Evaluation Results
Architectural Profile & Capabilities
Claude Haiku 5.5 is the fastest and most cost-efficient member of Anthropic's Claude 5.5 family, released October 7, 2026 and available on the Claude Platform, Claude Code, Amazon Bedrock/AWS, Google Cloud, and Microsoft Foundry. It succeeds Claude Haiku 4.5 (Oct 15, 2025), and Anthropic positions it for real-time, high-volume and cost-sensitive work: summarization, classification, routing, context compaction, subagents, computer/browser use, and coding sub-tasks. The headline engineering changes are a jump to a 1,000,000-token context window (up from the 200K tier) with a 128,000-token maximum output, plus a tiered pricing scheme that is unusually cheap for a long-context model: $0.10 per million input / $0.50 per million output tokens for prompts up to 100K tokens, escalating to $0.50 / $2.50 per million above 100K input. Anthropic also cites up to ~90% savings via prompt caching and ~50% via batch processing. Like its predecessor, Haiku 5.5 is a hybrid reasoning model (non-reasoning and extended-thinking modes); Anthropic has not publicly disclosed parameter count, layer count, hidden size, attention variant (GQA/MLA), or training-corpus details, so dense-vs-sparse cannot be confirmed from primary sources. Anthropic's own framing implies continued improvement in coding and computer use over Haiku 4.5, though the numeric system-card scorecard values for SWE-bench Verified, GPQA Diamond, MATH-500 and LiveCodeBench were not exposed in the retrievable excerpts. The model accepts text, images and PDFs and supports tool use and structured output. In the modern frontier context (GPT-6 Astra / GPT-5.6 Sol, Claude Opus 5 / Sonnet 5, Gemini 3.8 Flash, DeepSeek V4 Pro, Grok 4.6, Qwen 3.8, Kimi K3), Haiku 5.5 competes not on peak reasoning but on the price/latency/context Pareto frontier for fleet-scale orchestration.
Anthropic describes Haiku 5.5 as a hybrid-reasoning large language model in its small/fast tier — i.e., a dense autoregressive Transformer with an optional extended-thinking mode, not a sparse MoE, diffusion, or dLLM design. Anthropic does not publish active/total parameter counts, layer depth, hidden dimension, tokenizer, or the specific attention mechanism (e.g., GQA/MLA), so these cannot be confirmed from official sources. Context window is 1,000,000 tokens (system prompt + messages + tool definitions + tool results + images + documents) with up to 128,000 output tokens per request; test-time reasoning tokens are supported via the thinking budget. Hosted deployments are on Anthropic's own infrastructure plus AWS, Google Cloud, and Microsoft Foundry.
Recommended Workloads & Primary Use Cases
- •Tiered pricing cliff at 100K tokens: input jumps 5x ($0.10→$0.50/M) and output 5x ($0.50→$2.50/M) once a prompt exceeds 100K tokens, so treat 100K as a hard architectural budget boundary rather than a soft hint.
- •Extended-thinking mode inflates TTFT and billed output tokens; latency- and cost-sensitive routes should default to non-reasoning mode and reserve reasoning for genuinely hard steps.
- •Peak reasoning and long-horizon agentic accuracy still lag Claude Opus 5 / Sonnet 5 tiers — do not use Haiku 5.5 as the terminal decision-maker for complex, multi-step planning where a larger model materially outperforms it.
- •Official numeric scorecards for SWE-bench Verified, GPQA Diamond, MATH-500 and LiveCodeBench were not disclosed in retrievable primary sources; vendor-derived figures for the predecessor Haiku 4.5 vary widely (e.g., 73.3% vs 40.6% on SWE-bench Verified) depending on scaffold, thinking budget and subset, so cross-model comparisons must hold methodology constant.
- •Prompt-formatting sensitivity around tool definitions, structured output schemas and image/PDF grounding can degrade reliability at extreme context lengths; validate before assuming uniform quality across the full 1M window.
- •Anthropic does not publish parameter counts or attention architecture, limiting static capacity planning and self-hosting; fine-grained throughput/quantization control depends on the host provider.
Calls to Haiku 5.5 are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Anthropic.
More Models from Anthropic
Compare other engines available in this laboratory.
Claude Opus 5.5 is Anthropic's September 2026 flagship Opus-class model for long-horizon agentic coding, computer use and knowledge work — 1M-token context, native image input, adaptive (non-disableable) reasoning at four effort tiers, and $4/$20 per 1M tokens (40% cheaper per typical task than Opus 5). Available with Zero Data Retention inference on ARMES AI.
Anthropic's flagship 1M-context frontier multimodal foundation model, engineered for deep reasoning, autonomous coding agents, and complex long-horizon system execution.
Anthropic's most capable model (released May 28, 2026; knowledge cutoff Jan 2026).