Claude Haiku 5.5 launched October 7 at $0.10/$0.50 per million tokens, 90% below Haiku 4.5 on prompts under 100K. What the 5x long-prompt tier, the benchmarks and the Sonnet 5.5 cache cut mean for builders.

Anthropic released Claude Haiku 5.5 on October 7, calling it "the cheapest, fastest, and most capable small model we've ever released." It is the third and last model of the Claude 5.5 generation, arriving fifteen days after Opus 5.5 (September 22) and nine days after Sonnet 5.5 (September 28).
The headline is the price. On prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Haiku 4.5, the model it replaces, costs $1 and $5. That is a tenth of the price per token.
On the same day Anthropic halved the price of prompt-cache reads on Sonnet 5.5, and announced monthly API credits for Max and Team subscribers. For anyone running agents or high-volume pipelines on Claude, all three changes land on the same bill.
Every Claude model from 4.6 on bills its full 1M-token context at one flat rate. Haiku 5.5 is the exception: a request whose prompt runs past 100,000 tokens pays five times the short-prompt rate on everything, output included. Figures below were checked on Anthropic's pricing page on October 9, 2026, in dollars per million tokens.
| Model | Input | Output | Cache hit | Batch (in / out) |
|---|---|---|---|---|
| Haiku 5.5, prompt up to 100K | $0.10 | $0.50 | $0.01 | $0.05 / $0.25 |
| Haiku 5.5, prompt over 100K | $0.50 | $2.50 | $0.05 | $0.25 / $1.25 |
| Haiku 4.5 | $1 | $5 | $0.10 | $0.50 / $2.50 |
| Sonnet 5.5 | $2 | $10 | $0.10 (was $0.20) | - |
| Opus 5.5 | $4 | $20 | $0.20 | - |
The rule that matters is in the fine print. Anthropic's docs say a request's prompt length "counts all of its input tokens, including cache reads and cache writes," and that "a request over the threshold pays the higher prices even when part of its prompt is a cache hit." Caching makes a long context cheaper to resend. It does not move you back under the line.
A worked example shows the cliff. A request with a 90,000-token prompt and 2,000 tokens of output costs about $0.010. Grow the prompt to 150,000 tokens and it costs about $0.08 - eight times as much for a prompt two-thirds longer. That long request is still half what Haiku 4.5 would charge ($0.16), so nobody pays more than before. But the 10x saving only holds below 100K.
This is also why Anthropic's own savings claim is smaller than the table suggests. It says Haiku 5.5 "costs around 75% less to run" on average, a blended figure that includes long prompts. The per-token cut on short prompts is 90%. SiliconANGLE reports that about 90% of requests fit under 100,000 tokens, which is the bet behind the pricing.
Anthropic positions Haiku 5.5 against OpenAI's small model, GPT-6 Luna. On Anthropic's published benchmarks, Haiku 5.5 beats Luna on all six tests where both have a score. Luna has no published result on either Humanity's Last Exam row. Sonnet 5.5 stays ahead of Haiku 5.5 on every row.

The gaps over Haiku 4.5 are large enough to change what a small model is used for. Computer use on the OSWorld 2.1 offline subset went from 15.7% to 72.4%, against 48.9% for Luna. On Terminal-Bench 4.0, an agentic coding test, Haiku 4.5 scored 0.0% and Haiku 5.5 scores 39.2%. Luna gets 16.4%; Sonnet 5.5 gets 70.6%. Knowledge-work Elo on GDPval-AA v2.1 more than doubled, from 735 to 1620.
Coding is where the distance to Sonnet still shows. Terminal-Bench is 39.2% against 70.6%, so a Haiku-only coding agent will finish far fewer tasks. FrontierCode is closer, 46.4% against Sonnet's 52.1%, though Sonnet was run at its "Xhigh" effort setting there.
Haiku 5.5 is also the first Haiku with an adjustable effort setting, so a team can trade cost against quality per request instead of switching models. It has a 1M-token context window and up to 128,000 output tokens. Customer quotes in the launch post include Asana reporting over 30% lower latency and Box reporting scores 11 points above Haiku 4.5 at roughly half the latency.
The second change is quieter but may matter more to agent builders. Sonnet 5.5 prompt-cache reads dropped from $0.20 to $0.10 per million tokens, or 0.05x the base input price instead of 0.1x. According to the platform release notes, "cache writes and all other prices are unchanged."

Agentic loops resend the same long context on every turn, so cache reads are often the largest line on the bill. Anthropic estimates the cut makes Sonnet 5.5 "around 20% cheaper on most agentic work." Unlike Haiku 5.5, Sonnet 5.5 keeps flat pricing across its whole context, so a 400K-token cached session gets the full discount.
Anthropic also added monthly API credits for app subscribers: $100 a month on Max 5x, $200 on Max 20x and up to $500 on Team, pooled across users. They are claimed by linking a Claude Console organization.
The clearest winners are workloads that are short, repetitive and numerous: classification, summaries, extraction, support replies, database queries and subagents fanned out by a larger model. Anthropic names live customer support and browser use as the speed-sensitive cases. Moving those from Haiku 4.5 costs nothing but a model ID change, claude-haiku-5-5, and cuts per-token cost by 90%.

Long-context jobs need a closer look. A pipeline that feeds whole codebases or 200-page contracts into a small model will pay $0.50 / $2.50 per request on Haiku 5.5. That is still cheaper than Haiku 4.5, but the gap to Sonnet 5.5 narrows to 4x on input. If a job routinely sits just above 100K, trimming or compacting the prompt to get under the line is worth real money.
Who should not switch to Haiku alone: teams running serious coding agents. The Terminal-Bench gap says Sonnet 5.5 or a larger Claude model remains the right main model there, with Haiku 5.5 doing the cheap lookups underneath.
Haiku 5.5 is live in the Claude apps for Free, Pro, Max, Team and Enterprise users on web, iOS and Android, in Claude Code, on the Claude Platform, and through Amazon Web Services, Google Cloud and Microsoft Foundry. Anthropic published a system card dated October 7, reporting improvements on almost all of its alignment evaluations. Its cybersecurity safeguards are stricter than Haiku 4.5's, though somewhat looser than other recent Claude models'.

No retirement date for Haiku 4.5 has been announced. The one scheduled deprecation is Sonnet 4.5, which leaves the API on November 30, 2026; anyone still on it can compare against our coverage of Claude Sonnet 5. With Haiku 5.5 out, all three 5.5 models have shipped in under three weeks.
