Claude Haiku 5.5: Anthropic Cuts Small-Model Price by Up to 90%
Claude Haiku 5.5 is Anthropic's cheapest and fastest small model yet: $0.10 per million input tokens for prompts up to 100K, and the first Haiku with an effort setting.

- Anthropic says it is 90% cheaper than Haiku 4.5 for prompts up to 100,000 tokens, and about 75% cheaper to run on average.
- The docs list a 1 million token context window and up to 128,000 output tokens. It ships on the Claude Platform as claude-haiku-5-5, plus AWS, Google Cloud and Microsoft Azure.
- Alongside it, Sonnet 5.5 cache reads drop by half to $0.10 per million tokens, and Max and Team subscribers get monthly API credits.
in this block
Claude Haiku 5.5 is Anthropic's new budget workhorse. Released on October 7, it is the company's cheapest and fastest small model so far, priced at $0.10 per million input tokens for prompts up to 100,000 tokens, and it is the first Haiku with an adjustable effort setting.
What actually happened
Anthropic pitches the model for high-volume, cost-sensitive work: summaries, compaction, database queries, classification, and subagent jobs under a bigger model. Because it is also the fastest Claude at standard speed, Anthropic recommends it for live customer support and browser use.
The pricing table is the headline. For prompts up to 100,000 tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Above that, it is $0.50 and $2.50. Cache reads cost $0.01 or $0.05, and cache writes $0.125 or $0.625. SiliconANGLE notes Haiku 4.5 costs $1 input and $5 output.
Anthropic says about 90% of requests to the old Haiku fell under 100,000 tokens, which is why it leans on the 90% figure. Its 75% average also factors in a new tokenizer, so token counts for the same work changed too.
The benchmarks
On Anthropic's own chart, Claude Haiku 5.5 scores 72.4% on the OSWorld 2.1 offline subset for computer use. That compares with 15.7% for Haiku 4.5, 48.9% for GPT-6 Luna and 83.9% for Sonnet 5.5. On GDPval-AA v2.1, a knowledge-work test from Artificial Analysis, it posts 1620 versus 735 for Haiku 4.5 and 1840 for Sonnet 5.5.
Anthropic is upfront about limits. It says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding like Terminal-Bench 4.0, and that Haiku is best for narrower tasks that used to be too expensive to run at scale.
Customer quotes back up the speed pitch, with the usual launch-day caveat. SiliconANGLE highlights Asana, which says task completion latency fell more than 30% versus its current model, with inference per agent turn up to 2.5 times faster.
Safety and tooling
Anthropic says alignment tests found far fewer instances of misaligned behavior than with Haiku 4.5. Its cyber safeguards allow a wider range of defensive work than Sonnet 5.5's, but still block penetration testing and other attacker-style techniques, as SiliconANGLE also reports. Teams that need more can apply to Anthropic's Cyber Verification Program.
Anthropic also updated its Python and TypeScript SDKs with computer use and browser use in beta, and says the small model is especially suited to those tasks because of its speed and price.
The other price cuts
Sonnet 5.5 cache reads fall from $0.20 to $0.10 per million tokens. Anthropic says that makes Sonnet 5.5 about 20% cheaper on most agentic work, since cache reads make up a big share of tokens.
The new monthly API credits arrive this week: $100 for Max 5x, $200 for Max 20x and up to $500 pooled for Team plans, per Anthropic and SiliconANGLE. Credits can be used on any Claude model. On AWS, Bedrock offers the model through US, EU, Australia, Japan and global inference profiles, plus GovCloud.
What Polymarket thinks
Polymarket has a "Next Claude Haiku Model (4.6+): Text Arena Debut?" market, and Haiku 5.5 fits its naming rule. At 06:10 UTC on October 8, the Polymarket Haiku arena market had 1400+ at 90%, 1420+ at 79.5%, 1440+ at 40.5% and 1460+ at 19.5%. The 1420+ line jumped 14.5 points in a day.
It resolves on the model's Arena text score at noon ET on the day after it first appears on the leaderboard. Volume is small, about $20,800, so a single trader can move it.
What to do as a reader
If you run high-volume AI jobs, test Claude Haiku 5.5 against your current small model on real traffic. Classification, routing, summaries and subagent calls are where the price cut matters most.
Watch prompt length, because pricing jumps above 100,000 tokens, and keep a bigger model for hard coding work. If you have a Max or Team plan, use the new credits to try it for free. For context, see our Claude Sonnet 5.5 and Claude Opus 5.5 coverage, plus our Claude for Google Workspace explainer.
Not financial advice. DYOR, ser.