Signals Oct 10, 2026 5 min read FacebookX (Twitter)WhatsAppLinkedInPinterest

Anthropic Just Cut Haiku’s Price by 90 Percent. Here Is Why It Matters

Abstract black space-dust illustration

Anthropic released Claude Haiku 5.5 on October 7, and the headline number is hard to miss: 90 percent off. For prompts up to 100,000 tokens, input now costs $0.10 per million tokens and output $0.50, down from $1 and $5 on Haiku 4.5. That is the steepest price cut Anthropic has ever shipped, and it lands on its fastest model line, not its cheapest-to-run one. This is not a discount on old hardware. It is a bid for the middle of every AI product budget.

What actually changed

Haiku 5.5 is the third model in Anthropic’s Claude 5.5 family, following Opus 5.5 and Sonnet 5.5. It is built for high-volume, speed-sensitive work: summarization, classification, customer support replies, background agents, and subagent tasks where a larger model plans and Haiku executes.

Three things are new. First, the price: the two-tier structure keeps the headline rate only for prompts under 100,000 tokens; above that threshold, input rises to $0.50 and output to $2.50 per million tokens. Anthropic says about 90 percent of Haiku 4.5 requests stayed under that line, and after adjusting for a new tokenizer that counts roughly 30 percent more tokens for the same text, it estimates Haiku 5.5 runs about 75 percent cheaper on average.

Second, effort controls: Haiku 5.5 is the first Haiku model that lets developers dial how much thinking the model does on a request, from low to max. A simple email label can get low effort and low cost; a tricky classification can get more. Until now, that trade-off meant switching models entirely.

Third, capacity: the context window jumps from 200,000 tokens on Haiku 4.5 to one million, with up to 128,000 output tokens and a June 2026 knowledge cutoff. It is live on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure.

Anthropic also sweetened the rest of the lineup: Sonnet 5.5 cache reads were cut in half to $0.10 per million tokens, and Max and Team subscribers get monthly API credits of $100 to $500 depending on plan.

The real story is the price war, not the model

The most telling number in the announcement is not the 90 percent cut. It is $0.10 and $0.50, identical to OpenAI’s GPT-6 Luna rates at the same tier. Anthropic matched OpenAI’s small-model price exactly. That is deliberate. The small-model market has converged on a single list price, which means the buying decision moves away from cost per token and toward reliability, controls, and how well the model fits a builder’s stack.

This is part of a broader squeeze. Google just restricted free Gemini access this same week, pushing users toward paid tiers. Model access is getting more expensive for casual users at the same time it gets cheaper for serious builders. The industry is segmenting: free users pay in limits, heavy users pay in volume, and everyone in the middle is being pushed to pick a side.

The effort dial is the bigger idea

The price cut gets the headlines, but the adjustable effort setting is the more interesting engineering choice. In real products, not every request deserves the same spend. A background agent that summarizes thousands of documents a day should not burn premium reasoning on every file, but it should spend more on the ambiguous ones. Haiku 5.5 lets developers tune that inside one model instead of routing between three.

This matters most for the agent architectures Anthropic keeps pushing: subagents doing fast, repetitive tasks under a planner model’s direction. When every subagent call costs pennies instead of dollars, agent workflows that were previously too expensive to run at scale become viable. That is exactly the workload where Haiku 5.5 wants to live.

The catches nobody should ignore

First, the 100,000-token threshold. The $0.10 rate is a cliff: a 150,000-token prompt costs five times more per token than a 99,000-token one. Developers feeding long documents or large conversation histories into the model need to budget around that edge, and OpenAI’s Luna is cheaper on list price at that size because its higher tier only kicks in above 272,000 input tokens.

Second, the tokenizer. The same text counts as roughly 30 percent more tokens on Haiku 5.5 than on Haiku 4.5, which eats into some of the headline savings. Anthropic folded that into its 75 percent estimate, but individual bills will vary with prompt length, caching, and the effort setting.

Third, the benchmarks are vendor-reported. Anthropic claims large gains over Haiku 4.5, including a computer-use test jumping from 15.7 to 72.4 percent. Artificial Analysis ran its own index and ranked Haiku 5.5 at max effort second of 182 models in its class, but vendor numbers are a starting point for testing, not a result. Also note that non-default temperature, top p, and top k values now return a 400 error, so existing Haiku integrations need a migration pass, not just a price update.

What builders should do

Start with three moves. First, re-price your agent loops. If you run summaries, classifications, or subagents at volume, Haiku 5.5 is worth a pilot this week, with batch processing taking another 50 percent off input and output. Second, test the effort dial against your real tasks, because low effort only saves money if accuracy holds where it counts. Third, treat the model as a replaceable part: with small-model prices converging across vendors, switching costs matter more than list prices. Design your prompts and evaluation around that.

The broader lesson is the one from Google’s free-tier squeeze this week: the model market is splitting. Cheap, fast models for builders; tight limits for free users. The builders who win will be the ones who budget per task, test on their own data, and never let a single model’s pricing page become load-bearing architecture.

Sources

  • Anthropic announcement coverage: TechBooky, Particle, Startup Fortune, MarkTechPost, AppStack Insider, Breakread, SiliconAINews (all reporting the October 7, 2026 launch)
  • Pricing figures from Anthropic’s published pricing and platform documentation as reported by those outlets

Related reading

Add to preferred sources

Building something similar?

I write about the stack, the failures and the wins — follow along at the Lab.

Leave a Reply

Your email address will not be published. Required fields are marked *