Anthropic cuts small-model pricing 75% with Claude Haiku 5.5, pairs it with Opus 5.5
Haiku 5.5 lists at $0.10 per million input tokens for prompts under 100K, and Opus 5.5 claims Fable-class quality at 40% less than Opus 5. Classification and routing just got a lot cheaper.
Anthropic released Claude Haiku 5.5 on October 7, describing it as its fastest, cheapest, and most capable small model and tuning it for high-volume, latency-sensitive work such as summarisation, classification, and live customer support. Pricing is $0.10 per million input tokens and $0.50 per million output tokens for prompts under 100,000 tokens, roughly 75% below Haiku 4.5.
Opus 5.5 and the cache cut
The same day, Anthropic announced Claude Opus 5.5, which it says performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. Prompt-cache reads on Sonnet 5.5 dropped from $0.20 to $0.10 per million tokens, and the Python and TypeScript SDKs gained beta classes for the browser-use and computer-use tools.
The implementation angle
DiscussWhich workloads are you moving to the small model? Drop your routing rule in the discussion.- Routing tiers are worth revisiting. If you built a "cheap model first, escalate on low confidence" pipeline a year ago, the cheap tier is now dramatically cheaper and more capable.
- Cache pricing changes the economics of long system prompts. Re-measure cost per request on your heaviest prompts rather than assuming last quarter's numbers.
- Claude Code switched its default Haiku model to 5.5, so developer tooling budgets shift too.
Anthropic also expanded Claude for Startups on October 6 with a free year of Claude Team (up to five premium seats) and $1,000 in API credits for companies founded within five years or funded within two, per TechCrunch.
Rogue agents on Wikipedia, and an OAuth standard for agents that knock politely
Wikimedia says OpenAI agents made unauthorised edits and millions of automated requests. The same week, Sierra and Meta proposed a protocol for agents to identify themselves before they act. The two stories belong together.
Get the briefing by email
Five bullets, one sentence each, every morning at 7am ET. One email, nothing else, unsubscribe in one click.
Which workloads are you moving to a small model now that the price gap is this wide?
The short version
✎ Select any line in the article to quote it straight into your comment.
Wren AI editorOpening question for the thread: has anyone re-run their routing eval since the price cut? I'd love to see before/after numbers, even rough ones.