Analysis
Claude Haiku 5.5: what the new price means for enterprise AI costs
1AYM is an OpenAI Select Partner.
Published . Vendor facts checked on against each vendor’s own pages, listed at the end. Plans, prices and availability change. Neither OpenAI nor Anthropic has reviewed this analysis.
The short version
One release, three things to act on: the price, the token count and the code changes. Each has its own section, with what to check before moving a workload. My take, from testing it beside Opus 5.5, is in the last section.
- The price and the 100,000-token step
- $0.10 in and $0.50 out per million tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 above. Anthropic says about 90% of Haiku 4.5 requests were under the line.
- About 30% more tokens for the same text
- Haiku 5.5 uses Anthropic's newer tokenizer, so the same text is about 30% more tokens than on Haiku 4.5. Budgets, limits and cost estimates measured in tokens all move.
- Breaking changes from Haiku 4.5
- Five breaking changes, including errors on manual thinking budgets, temperature settings and assistant prefill. Changing the model name is not the whole migration.
The price, and the step at 100,000 tokens
What was announced
Claude Haiku 5.5 was released on 7 October 2026 and is available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic describes it as built for high-volume, latency-sensitive work such as classification, routing, extraction and subagent tasks, and as its fastest model at standard speed.
For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens. For prompts over 100,000 tokens the rates are $0.50 and $2.50. Cache reads are $0.01 and $0.05 on the same two bands, and the Batch API takes 50% off input and output. Haiku 4.5 is $1.00 in and $5.00 out.
Anthropic says Haiku 5.5 costs around 75% less to run than Haiku 4.5 on average. Its footnote explains the sum: the price is 90% lower for requests up to 100,000 tokens and 50% lower above, 90% of Haiku 4.5 requests fell in the lower band, and the figure allows for the new model using more tokens for the same work. The context window is 1M tokens with up to 128k of output, up from 200k and 64k.
What it means for enterprises
The work this model is priced for is the dull, high-volume part of an AI estate: classifying tickets, routing requests, pulling fields out of documents and running as a subagent under a larger model. Those jobs rarely make a board paper, but in a production estate they can be most of the token bill, and a price a tenth of the old one changes which of them are worth automating at all.
The 75% figure is an average over Anthropic's own traffic. Your number depends on where your prompts sit against the 100,000-token line. A classification job with a short prompt lands in the cheap band every time. A job that loads long contracts, a whole case history or a large codebase into the 1M window pays the higher rates, and the larger window makes it easier to cross the line without noticing.
Being on all three large clouds from day one matters for procurement. An organisation standardised on AWS, Azure or Google Cloud can usually buy the model through the agreement it already has.
Before you switch it on
- Prompt lengths
- Pull a month of request logs and count how many prompts are over 100,000 tokens once recounted for the new tokenizer. That share sets your real saving.
- Cost
- Model the bill on both bands, with caching and batch discounts where you use them, before quoting 75% to a finance team.
- Quality
- Run your own evaluation set. Anthropic's benchmark results are its own, and a cheaper model that needs retries is not cheaper.
- Procurement
- Check availability in the cloud region your data has to stay in. The launch pages name the platforms, not the regions.
The same text now counts as about 30% more tokens
What was announced
Haiku 5.5 uses the same newer tokenizer as Claude 4.7 and later models. Anthropic's documentation says the same input text produces approximately 30% more tokens on Haiku 5.5 than on Haiku 4.5, and that the exact increase depends on the content. The shape of requests and responses does not change. Anything measured or budgeted in tokens does.
What it means for enterprises
A per-token price cut and a per-text price cut are different numbers. On Anthropic's published rates, a prompt that stays under 100,000 tokens costs about 87% less for the same text, since 30% more tokens at a tenth of the price is 13% of the old cost. Above the line it is about 35% less, since 30% more tokens at half the price is 65% of the old cost. Both figures are our arithmetic on the published rates, and the second is a long way from 75%.
The tokenizer also moves the line itself. A prompt that measured 80,000 tokens on Haiku 4.5 is roughly 104,000 on Haiku 5.5, which puts it in the higher band. Any workload sitting between about 77,000 and 100,000 tokens today needs checking first, because it changes band without anyone changing the prompt.
Token-based controls move too. Rate limits, max_tokens settings, per-team budgets and chargeback reports that were sized on Haiku 4.5 counts will all read about 30% high for the same work.
Before you switch it on
- Recount
- Recount your real prompts with the new tokenizer before estimating anything. Do not scale the old numbers by a flat 30%.
- The band edge
- List every workload whose prompts sit above about 77,000 tokens on Haiku 4.5. Those are the ones that can cross 100,000.
- Limits
- Revisit max_tokens and rate limits. Thinking tokens count towards max_tokens, so a tight limit can stop a response before any text arrives.
- Reporting
- Tell whoever owns AI spend reporting that token volumes will rise while the bill falls, so the first monthly report is not read as a problem.
Breaking changes for code written for Haiku 4.5
What was announced
Anthropic lists five breaking changes from Haiku 4.5. Manual extended thinking with a token budget returns an error and is replaced by adaptive thinking. A non-default temperature, top_p or top_k returns a 400 error. Assistant message prefill returns an error, so a request has to end on a user turn. Computer use needs a new toolset on the Claude API and Google Cloud. Changing earlier turns of a conversation invalidates the thinking blocks sent back with it.
Behaviour changes as well. Adaptive thinking is on by default, so a response can begin with a thinking block, and code that reads the first content block as the answer will read the wrong thing. Thinking text is omitted unless you ask for a summary. Safety classifiers can decline a request, which arrives as a refusal stop reason your client has to handle.
Haiku 4.5 is still listed as active. Anthropic's deprecations page gives no retirement date for it, only that retirement will be not sooner than 15 October 2026, and it says customers get at least 60 days' notice. Haiku 5.5 is listed as not retiring before 7 October 2027. Those dates cover the Claude API, Claude Platform on AWS and Microsoft Foundry; Amazon Bedrock and Google Cloud set their own.
What it means for enterprises
Each of these is small on its own. Together they mean a production service will return errors the day someone changes the model name without reading the migration guide. Temperature and prefill are the two most likely to bite, because both were common ways to get short, predictable output from a small model.
The practical question is who owns the change. A team with one service and good tests can migrate in a few days. An estate where a dozen teams each call Haiku 4.5 in their own way needs a list of callers, an owner for each and an evaluation run before and after, or the saving arrives with a run of incidents.
No retirement date has been announced for Haiku 4.5, and Anthropic gives at least 60 days' notice of one, so the migration can follow your release process, with the highest-volume workload first because that is where the money is.
Before you switch it on
- Inventory
- Search your code for temperature, top_p, top_k, budget_tokens and prefilled assistant turns on Haiku calls.
- Response parsing
- Select content blocks by type, not by position, and handle a refusal stop reason.
- Evaluation
- Run the same test set on Haiku 4.5 and Haiku 5.5 before switching, and compare cost per completed task, not cost per token.
- Rollout
- Move one workload at a time behind a flag, and keep Haiku 4.5 available as the fallback until the numbers hold.
Where the market is heading: small models priced for agent work
Anthropic is explicit about the job. It positions Haiku 5.5 as a subagent beside its larger models, and says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. The small model takes the narrow, repeated steps that were too expensive to hand to a model before, such as compaction, summarisation and subagent work.
The same day, Anthropic halved the price of cache reads on Sonnet 5.5, from $0.20 to $0.10 per million tokens, and says that makes Sonnet 5.5 around 20% cheaper on most agentic tasks. Read together, the two changes cut the cost of long-running agents from both ends: the lead model re-reads its context for less, and the helpers under it cost a tenth of what they did.
Anthropic's own results put Haiku 5.5 at 72.4% on the OSWorld 2.1 computer-use benchmark, against 15.7% for Haiku 4.5, both on an offline subset. Those are the vendor's numbers. If they hold up in independent testing, operating software through its screens stops being a job only the expensive models can do.
For a buyer, the consequence is that model choice becomes routing. A serious agent system will run two or three models at different prices, and the saving comes from sending each step to the cheapest one that passes your evaluation.
My take
I have been testing Haiku 5.5, and it works well paired with Opus 5.5 for faster and cheaper workflows: Opus does the planning and the review, and Haiku does the execution only. I have barely seen a dent in my usage. Opus was already efficient, and this makes it even better.
Sources: [1]
What is not known yet
The figures above come from Anthropic's announcement and documentation, published the day before this post. Some questions they do not answer.
- Independent results
- The benchmark scores and the 75% average are Anthropic's. We have not seen independent tests of quality or cost.
- Regions and data residency
- The pages we read name the cloud platforms but not the regions, or the data residency options on each.
- When Haiku 4.5 retires
- No retirement date is published, only a not-sooner-than date of 15 October 2026.
- Your own average
- The 90% of requests under 100,000 tokens is Anthropic's traffic on Haiku 4.5. A document-heavy workload can look very different.
For engineers: model IDs, the migration list and effort settings
Model IDs and limits
The model ID is claude-haiku-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and anthropic.claude-haiku-5-5 on Amazon Bedrock. Context is 1M tokens and output is up to 128k, or 300k on the Message Batches API with the output-300k-2026-03-24 beta header. Input is text and images.
Sources: [2]
The migration list
Replace budget_tokens with adaptive thinking. Omit temperature, top_p and top_k. End messages on a user turn, since prefill returns an error. Replace computer_20250124 with computer_toolset_20260801 on the Claude API and Google Cloud. Keep conversations append-only if you send thinking blocks back, and replay stored conversations through the account that produced them.
Select content blocks by their type field, because a response can start with thinking blocks. Set thinking.display to summarized if you need the thinking text. Handle stop_reason refusal in the client, as server-side fallback is not available.
Sources: [3]
Effort
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, and the default is medium. Effort is the intended way to trade quality against speed and cost, so treat it as a parameter in your evaluation: run the test set at two or three effort levels and price each, since a lower setting on the cheap band may beat a larger model on cost per completed task.
Where 1AYM fits
A model change like this is part of running AI in production, which is what our Production AI systems work is for. The steps are the ones above: recount the tokens on your real prompts, run the evaluation on both models, migrate the code and put a budget on the result. The saving is only real once it shows up in a month of production spend.
If the question is still which workloads are worth moving, we would start with our AI Opportunity & Feasibility Sprint, typically two to four weeks, which ranks them by value, feasibility and risk before anyone changes a model name. It does not have to start big: 1AYM takes small fixed-scope statements of work as well as larger builds.
Sources
- [1]Anthropic, Introducing Claude Haiku 5.5, 7 October 2026 (checked 8 October 2026)
- [2]Anthropic, Claude Platform documentation: Claude Haiku 5.5 overview (checked 8 October 2026)
- [3]Anthropic, Claude Platform documentation: What's new in Claude Haiku 5.5 (checked 8 October 2026)
- [4]Anthropic, Claude Platform documentation: Model deprecations (checked 8 October 2026)
Frequently asked questions
How much does Claude Haiku 5.5 cost?
For prompts up to 100,000 tokens, $0.10 per million input tokens and $0.50 per million output tokens. For prompts over 100,000 tokens, $0.50 and $2.50. Cache reads cost $0.01 and $0.05 on the same bands, and the Batch API takes 50% off input and output. Prices are Anthropic's, checked on 8 October 2026.
Is Claude Haiku 5.5 really 75% cheaper than Haiku 4.5?
On average across Anthropic's own traffic, by its own calculation. The price is 90% lower for requests up to 100,000 tokens and 50% lower above that, and the same text counts as about 30% more tokens. By our arithmetic, a workload with short prompts saves more than 75% and one with long prompts saves much less.
Is moving from Haiku 4.5 to Haiku 5.5 only a change of model name?
No. Anthropic lists five breaking changes: manual thinking budgets, non-default temperature, top_p or top_k, and assistant prefill all return errors, computer use needs a new toolset on the Claude API and Google Cloud, and editing earlier turns invalidates thinking blocks. Responses can also begin with a thinking block.
Where is Claude Haiku 5.5 available?
On the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, from its release on 7 October 2026. The launch pages do not list regions, so check the one your data has to stay in.
Further
- AI spend governance · How seats, tokens and metered agent runs add up, and how to budget, cap and charge them back.
- AI implementation cost · The six lines of an AI budget and what moves each one, of which model usage is only one.
- AI agents for business · Where agents pay back, where they do not, and the controls to set before a cheap model runs thousands of steps.
- Private LLM · Vendor API, private cloud or self-hosted: the hosting routes to weigh against a cheaper hosted model.
Working out what you would save?
Half an hour is enough to look at your highest-volume workloads and tell you which are worth moving first.
Last reviewed · 1AYM