The launch in brief
On 7 October 2026 Anthropic released Claude Haiku 5.5, the small, fast tier of its Claude 5.5 family, and cut its list price by 90% versus Haiku 4.5 for requests under 100,000 tokens. Anthropic estimates real workloads cost about 75% less once request sizes and token consumption are factored in.
The model is live under the API name claude-haiku-5-5 on the Claude Platform, AWS, Google Cloud and Microsoft Azure. It completes the 5.5 refresh that began with Opus 5.5 on 22 September and Sonnet 5.5 on 28 September.
For teams already building on Anthropic’s Claude models, the practical effect is that work previously routed to Sonnet for quality reasons may now fit on Haiku, and high-volume pipelines that were too expensive to run on an LLM at all may now pencil out.
How much does Claude Haiku 5.5 cost?
Haiku 5.5 introduces a two-tier price based on prompt length. Requests up to 100,000 tokens cost $0.10 per million input tokens and $0.50 per million output tokens. Above that threshold, the rate rises to $0.50 and $2.50, which is still half of what Haiku 4.5 charged. Prompt caching follows the same split, with cache reads at $0.01 per million tokens in the lower tier.
VentureBeat reports that roughly 90% of requests fall into the cheaper tier. That matches what we see in production: most classification, extraction and routing calls are short, while long-context calls cluster in a few document-heavy flows. The base rate exactly matches OpenAI’s GPT-6 Luna; Luna’s long-context surcharge starts later, at 272,000 tokens, which matters if your pipeline regularly sends large documents.
Anthropic also halved Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens, which SiliconANGLE estimates trims about 20% off typical agentic workloads that re-read the same system prompt and tool definitions on every turn.
How much better is it than Haiku 4.5?
On Anthropic’s published results the gap is large, particularly for agentic tasks. Haiku 5.5 scores 72.4% on OSWorld 2.1 computer-use tasks versus 15.7% for Haiku 4.5, and 39.2% on Terminal-Bench 4.0, where its predecessor scored zero. On the GDPval-AA knowledge-work benchmark it rates 1,620 against 735.
Customer reports point the same way. Asana measured more than 30% lower task latency and up to 2.5× faster inference; Box reported scores 11 points above Haiku 4.5 at about half the latency. The model also gains the adjustable effort setting from the larger 5.5 models, so teams can trade latency for accuracy per call.
Treat these numbers as a starting hypothesis. They are vendor-reported, several depend on the effort level used, and Anthropic has not published throughput in tokens per second. Your own evaluation set is the only benchmark that counts.
Why is Anthropic cutting prices now?
Because small-model pricing has become a competitive front. Yahoo Finance framed the launch as part of an intensifying AI pricing war: enterprise buyers are pushing for lower inference bills and increasingly test open-weight models from US and Chinese labs as alternatives. Matching GPT-6 Luna’s base price removes the easiest argument for switching vendors on cost alone.
The timing also fits Anthropic’s run-up to an expected IPO. Three model refreshes in 15 days, each with a lower price than its predecessor, signal that the company wants to win volume workloads, not only premium reasoning traffic.
What it means for US & EU software teams
Model routing deserves a fresh look. Many production systems send everything to a mid-tier model because the small tier failed on edge cases. With a 90% list-price cut and much stronger tool use, a router that sends easy calls to Haiku and escalates hard ones to Sonnet or Opus can cut LLM spend sharply without a visible quality drop. The saving only materialises if you measure escalation rates and failure modes per task type.
New use cases cross the cost line. At $0.10 per million input tokens, tagging every support ticket, summarising every call transcript or pre-screening every document upload becomes cheap enough to run by default. Subagent patterns, where one orchestrator fans out dozens of small calls, also become far cheaper.
The 100K threshold is an architecture input. Prompts above 100,000 tokens cost five times more per token. Retrieval that sends only relevant chunks, plus aggressive prompt caching of stable context, keeps most calls in the cheap tier. Dumping whole repositories or contract bundles into the context window now has a clear price signal against it.
Procurement and residency stay the same. Because Haiku 5.5 ships on AWS, Google Cloud and Azure on day one, EU teams can keep using their existing cloud region, data processing agreement and billing. Check that the specific region you rely on lists the model before planning a migration, and keep model-version pinning in place so a silent alias change does not shift behaviour under GDPR-relevant workflows.
How to evaluate a switch
- Pull a week of real traffic per task type and replay it against Haiku 5.5 at two effort levels, scoring against your existing golden answers.
- Measure the token mix. Calculate what share of calls exceeds 100,000 tokens; those are the ones to fix with retrieval or caching first.
- Add a confidence-based escalation path to Sonnet or Opus instead of a hard switch, and log every escalation.
- Re-check tool-calling and structured output on your own schemas; benchmark gains do not guarantee identical JSON behaviour.
- Pin the model version in configuration and roll out behind a feature flag with a cost and quality dashboard.
Frequently asked questions
How much does Claude Haiku 5.5 cost?
For prompts up to 100,000 tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 tokens the price is $0.50 input and $2.50 output. Cache reads cost $0.01 per million tokens in the lower tier. Haiku 4.5 cost $1 input and $5 output.
Where is Claude Haiku 5.5 available?
Anthropic released Haiku 5.5 on 7 October 2026 on the Claude Platform under the model ID claude-haiku-5-5, and on Amazon Web Services, Google Cloud and Microsoft Azure. Teams should confirm availability in the specific cloud region they use.
Is Claude Haiku 5.5 better than Haiku 4.5?
On Anthropic’s published benchmarks, yes, by a wide margin on agentic tasks: 72.4% vs 15.7% on OSWorld 2.1 and 39.2% vs 0% on Terminal-Bench 4.0. These figures are vendor-reported, so teams should validate on their own evaluation sets before switching.
Should we move production workloads from Sonnet to Haiku 5.5?
Not wholesale. The safer approach is a router that sends simple, high-volume calls such as classification, extraction and summarisation to Haiku 5.5 and escalates difficult requests to Sonnet or Opus, with quality and cost tracked per task type.
Sources
Anthropic — Introducing Claude Haiku 5.5 (7 October 2026)
CNBC — Anthropic unveils a new, cheaper Haiku model (7 October 2026)
Yahoo Finance — Anthropic reveals Haiku 5.5 model as AI pricing war intensifies (7 October 2026)
VentureBeat — Anthropic launches Claude Haiku 5.5 with 90% API price reduction, matching GPT-6 Luna (7 October 2026)
SiliconANGLE — Anthropic releases Claude Haiku 5.5 small model and halves Sonnet 5.5 cache read prices (7 October 2026)