Anthropic Launches Claude Haiku 5.5 for Cheaper, Faster Agent Work
Anthropic has introduced Claude Haiku 5.5, a smaller model built for high-volume work at lower cost and lower latency. The release also sharpens the economics of agentic workflows across major cloud platforms.
Anthropic puts a cheaper small model into production
Anthropic introduced Claude Haiku 5.5 on October 7, 2026, and described it as its cheapest, fastest, and most capable small model to date. That matters because many production AI systems do not need a large model for every step. They need a fast, inexpensive model for the bulk of routine work, with larger models reserved for harder decisions.
The model is aimed at high-volume, cost-sensitive tasks such as summaries, compactions, database queries, and classification requests. Anthropic also says it is suited for subagent roles in coding workflows and for speed-sensitive use cases like live customer support and browser use. For product teams, that makes Haiku 5.5 relevant not as a flagship research system, but as a practical building block for deployed agent workflows.
Performance gains are biggest where small models usually lag
Anthropic’s benchmark table shows large jumps over Haiku 4.5 in areas that matter for agentic systems. On OSWorld 2.1 offline subset, Haiku 5.5 scored 72.4%, compared with 15.7% for Haiku 4.5. On Humanity’s Last Exam, it reached 45.9% without tools and 57.4% with tools, versus 10.2% and 18.7% for Haiku 4.5.
The model also posted 39.2% on Terminal-Bench 4.0, while Haiku 4.5 scored 0.0%. Anthropic says Haiku 5.5 is its first Haiku-class model with an adjustable effort setting, which lets users trade cost against intelligence. That combination is important because it gives teams more control over where they want to spend tokens and latency, instead of treating every task the same way.
Pricing and platform access make the release immediately usable
The pricing table puts Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. For prompts over 100,000 tokens, pricing rises to $0.50 per million input tokens and $2.50 per million output tokens. Those numbers position the model for workloads where throughput and cost discipline matter as much as capability.
Anthropic says Haiku 5.5 is available now across AWS, Google Cloud, and Microsoft Azure, and that developers can use the model ID claude-haiku-5-5 on the Claude Platform. That broad availability is part of the story: a cheaper model only changes product economics if teams can actually deploy it in the environments they already use.
Anthropic also lowers the cost of its larger agent workflows
Alongside the Haiku launch, Anthropic cut Claude Sonnet 5.5 cache-read pricing by 50%, from $0.20 to $0.10 per million tokens. Anthropic says that change lowers Sonnet 5.5’s cost on most agentic tasks by around 20%. In practice, that means the company is pushing down costs at multiple points in the stack, not just at the small-model layer.
The broader message is that frontier competition is no longer only about raw intelligence. It is also about the economics of deployment: token prices, latency, routing, and how much of a workflow can be safely shifted to a smaller model. For teams building agents, the useful question is not whether Haiku 5.5 is a replacement for larger models, but where it can absorb routine work and reduce the number of expensive calls.
Safety boundaries still shape where teams can use it
Anthropic says Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, while still allowing a wider range of defensive tasks than Sonnet 5.5’s safeguards. It also says the model’s biology safeguards are the same as those for Sonnet 5, Sonnet 5.5, and Opus 5. For organizations working in these areas, the launch is not just about performance or price; it is also about what kinds of workflows remain available.
Anthropic points organizations to its Life Sciences Verification Program and Cyber Verification Program for broader work. For product and platform teams, the next step is to map the model against real workload tiers: use Haiku 5.5 for summaries, classification, and subagent execution; reserve larger models for harder reasoning; and watch how the new effort setting and lower prices affect routing and overall cost.
Sources
- Introducing Claude Haiku 5.5 Anthropic · October 7, 2026
- Newsroom Anthropic · October 8, 2026
- Google Cloud introduces the Gemini agent. Google Blog · October 8, 2026
- 2026 Usage Policy update Anthropic · October 8, 2026
This article was researched and drafted with AI from the sources above. Spot an error? Email info@hiddenproai.online.