Haiku 4.5
Anthropic's high-speed, frontier-class efficient model delivering 73.3% SWE-bench Verified coding performance at $1/$5 per million tokens with extended thinking capabilities.
Empirical Evaluation Results
Architectural Profile & Capabilities
Anthropic's Claude Haiku 4.5 represents the high-efficiency, speed-tier breakthrough in the Claude 4.5 generation. Engineered to match or surpass previous flagship intelligence (such as Sonnet 4) while delivering a ~2x to 3x reduction in latency and operational expenditure ($1.00 / 1M input tokens and $5.00 / 1M output tokens), Haiku 4.5 establishes a new frontier for agentic efficiency. The model features a 200,000-token context window with up to 64,000 output tokens and native multimodal vision capabilities. Crucially, Haiku 4.5 introduces support for test-time compute scaling ('extended thinking'), enabling the model to dynamically generate internal reasoning traces to achieve an exceptional 73.3% on SWE-bench Verified. In production architectures, Haiku 4.5 serves as the primary engine for high-throughput sub-agents, repository refactoring loops, classification routers, and low-latency interactive applications across Claude Code and enterprise developer tools.
Proprietary dense autoregressive multimodal transformer trained with Constitutional AI (RLAIF) and RLVR. Employs multi-query/grouped-query attention (GQA) optimized for low latency and high-throughput KV caching. Features dynamic test-time compute scaling (extended thinking mode with configurable token budgets) alongside native computer use, vision tokens, and prompt caching primitives (5-minute and 1-hour retention).
Recommended Workloads & Primary Use Cases
- •Extended thinking tokens consume the standard output window budget (up to 64,000 max tokens) and incur output pricing ($5/MTok).
- •While SWE-bench Verified (73.3%) rivals flagship frontier predecessors, deeply abstract architectural reasoning and advanced multi-step proofing lag behind Claude 4.5 Opus / Sonnet 4.5.
- •Time-to-first-token (TTFT) scales linearly on uncached long-context prompts; prompt caching headers must be actively managed to sustain sub-second latency across 200k context.
- •Computer-use capabilities are sensitive to high-DPI resolution scaling and require explicit downsampling coordinates compared to frontier desktop models.
Calls to Haiku 4.5 are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Anthropic.
More Models from Anthropic
Compare other engines available in this laboratory.
Anthropic's flagship 1M-context frontier multimodal foundation model, engineered for deep reasoning, autonomous coding agents, and complex long-horizon system execution.
Anthropic's most capable model (released May 28, 2026; knowledge cutoff Jan 2026).
Successor to Opus 4.6 at the same price (released Apr 16, 2026).