A lower token price helps only when the smaller model completes the real task reliably at production volume
Source: Anthropic
Anthropic says Haiku 5.5 is available now through its platform and on AWS, Google Cloud, and Microsoft Azure. It positions the model for high-volume, bounded tasks and adds an adjustable effort setting. On Anthropic’s platform, the listed rates are $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens; both rates are five times higher above that prompt length. Anthropic reports lower average cost than Haiku 4.5, while noting that tokenisation and task mix affect the realised saving.
Why this matters: Choose a repeatable task before changing model defaults: for example, classification, document compaction, or a narrow subagent step. Replay representative short and long inputs against the current model and Haiku 5.5, then record task success, error and escalation rates, latency, cache use, retries, and total cost per accepted result. Set quality thresholds and a fallback for difficult cases. Keep data access and tool permissions tied to the workflow, because a cheaper model does not make the surrounding agent safer or more accurate by itself.
Read Anthropic’s Haiku 5.5 announcement