The news
Anthropic released Claude Haiku 5.5, its newest small model, on October 7, 2026, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. On the same day it halved cache-read prices for Claude Sonnet 5.5, according to SiliconANGLE and Let's Data Science.
Those short-prompt rates are a tenth of what Haiku 4.5 cost, which was $1 per million input tokens and $5 per million output tokens, both outlets reported. Prompts above 100,000 tokens cost $0.50 for input and $2.50 for output per million tokens. Anthropic says that, on average, Haiku 5.5 costs about 75% less to run than Haiku 4.5. Let's Data Science reported that the lower short-prompt tier covers roughly 90% of requests sent to the previous Haiku model.
Haiku 5.5 is the first Haiku model with an adjustable effort setting, which lets developers trade token use against capability. Anthropic is aiming it at repetitive, high-volume tasks such as summaries, classification, coding subagents, customer support and browser automation. It is available through Anthropic's own platform as well as AWS, Google Cloud and Microsoft Azure, SiliconANGLE reported.
Anthropic reported that Haiku 5.5 scored 72.4% on the OSWorld 2.1 agent benchmark, against 15.7% for Haiku 4.5 on the offline subset, per Let's Data Science, and 48.9% for OpenAI's GPT-6 Luna, per SiliconANGLE. These are vendor-reported results. Asana, which tested the model, said task completion latency was more than 30% lower, SiliconANGLE reported.
For Sonnet 5.5, the price of reading cached tokens, meaning prompt content stored and reused across calls, fell from $0.20 to $0.10 per million tokens. Anthropic estimates this makes Sonnet 5.5 about 20% cheaper for most agentic work.
The numbers
- Haiku 5.5 input price (up to 100K tokens)
- $0.10 per million tokens
- Haiku 5.5 output price (up to 100K tokens)
- $0.50 per million tokens
- Haiku 4.5 prices
- $1 input / $5 output per million
- Average cost cut vs Haiku 4.5 (Anthropic)
- about 75%
- Sonnet 5.5 cache-read price
- $0.20 to $0.10 per million tokens
Why CEOs should care
For CFOs and heads of AI platforms, this is a direct cut to run costs on the highest-volume workloads: support chat, document triage, classification and the small subagents that sit under larger coding or research agents. Re-price your current Haiku 4.5 and Sonnet workloads at the new rates, and check how many of your prompts fall under the 100,000-token threshold where the lowest prices apply.
Engineering leaders should revisit their caching strategy. Agents tend to resend long instructions and context on each step, and cheaper cache reads reward teams that structure prompts so that shared content is cached. Ask your team what share of Sonnet spend is cache reads today; that is where the estimated 20% saving would come from.
Buyers should treat the benchmark claims with care. The OSWorld and latency numbers come from Anthropic and a customer, not independent tests. Run Haiku 5.5 on your own task samples, using the new effort setting, before moving traffic from larger models.
The bigger picture
Model makers are competing hard on price for small, fast models because agent systems call them many times per task. Each step in an agent workflow is a model call, so unit price drives the economics of automation at scale.
The release also comes as Anthropic prepares for a public listing. Cutting prices on widely used models can drive volume, but it also puts pressure on margins, a trade-off investors will watch as the company's financials become public.
What’s next
Expect rival vendors to answer with their own small-model price cuts. Customers on the Claude platform, AWS, Google Cloud or Azure can test Haiku 5.5 now and compare it with their current models on cost per completed task.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error
Companies in this story








