Anthropic released Claude Haiku 5.5 on October 7, describing it as its fastest, cheapest, and most capable small model and tuning it for high-volume, latency-sensitive work such as summarisation, classification, and live customer support. Pricing is $0.10 per million input tokens and $0.50 per million output tokens for prompts under 100,000 tokens, roughly 75% below Haiku 4.5.

Opus 5.5 and the cache cut

The same day, Anthropic announced Claude Opus 5.5, which it says performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. Prompt-cache reads on Sonnet 5.5 dropped from $0.20 to $0.10 per million tokens, and the Python and TypeScript SDKs gained beta classes for the browser-use and computer-use tools.

The implementation angle

DiscussWhich workloads are you moving to the small model? Drop your routing rule in the discussion.
  • Routing tiers are worth revisiting. If you built a "cheap model first, escalate on low confidence" pipeline a year ago, the cheap tier is now dramatically cheaper and more capable.
  • Cache pricing changes the economics of long system prompts. Re-measure cost per request on your heaviest prompts rather than assuming last quarter's numbers.
  • Claude Code switched its default Haiku model to 5.5, so developer tooling budgets shift too.

Anthropic also expanded Claude for Startups on October 6 with a free year of Claude Team (up to five premium seats) and $1,000 in API credits for companies founded within five years or funded within two, per TechCrunch.