← All reviews
Analysis · October 7, 2026 · 10 min read

Haiku 5.5's 90% price cut stops at 100K tokens — 95% of our Claude Code tokens were past it

Price the token counts Claude Code actually generated on Haiku 5.5's two sheets and the result is about half of what Haiku 4.5's prices give, not a tenth: 77.9% of calls and 94.8% of input-side tokens sat in prompts over 100K, where every rate is five times the headline. Subagent calls were over the line half the time. The model itself is the good news — it found all nine planted problems in our prompt audit, three runs out of three, where Haiku 4.5 had missed the broken commands in every run on 4 October.

The 30-second version

Anthropic released Claude Haiku 5.5 on 7 October. The headline prices are $0.10 per million input tokens and $0.50 per million output tokens, a tenth of Haiku 4.5, but they apply only to requests whose prompt is 100,000 tokens or shorter. Over that line, input is $0.50, output $2.50, and cache reads and writes rise by the same factor of five. Anthropic’s own summary is 90% cheaper below the line, 50% cheaper above it, and “around 75% less” on average.

We wanted to know which side of the line agent traffic sits on. So we took every Claude Code transcript on one of our machines — 208 transcripts, 12,901 distinct API calls between 10 July and 7 October — and measured the prompt size of each call:

  • 77.9% of calls had prompts over 100K. The median prompt was 209,571 tokens. In main sessions it was 286,661, and 90.2% of calls were over the line; in subagent transcripts the median was 101,891, so half of those calls were over too.
  • 94.8% of input-side tokens were on the expensive side. Priced on Haiku 5.5’s two-rate sheet, the recorded token counts come to $395.80, against $833.14 at Haiku 4.5’s rates: 52.5% less. At the headline rate on every call they would come to $83.31. Calls over 100K account for 98.7% of the Haiku 5.5 figure. These are modeled costs — the sessions ran on Opus and Fable, and over half of the calls were larger than Haiku 4.5’s 200K context window could have taken.
  • The same text is more tokens on Haiku 5.5. Anthropic says about 30% more. We measured 28% more for three source files and 40% more for a documentation page. Under a uniform 30% token-inflation assumption, the modeled saving on our traffic falls to 38.2%.
  • Claude Code 2.1.293 handles both. Its cost figures matched the over-100K sheet to the cent, and on the Anthropic API the haiku alias now resolves to Haiku 5.5 unless you override it, which moves /model haiku, model: haiku subagents and, when your model setting is haiku, sessions saved on Haiku 4.5 onto the new model.
  • The model is better. On the prompt-audit test we published on 4 October, Haiku 5.5 found all nine planted problems in three of three runs, listing no controls, for an average of $0.04 a run. Haiku 4.5, rerun the same day, found three, five and eight for $0.22.

Verdict: price Haiku 5.5 at the over-100K sheet for anything that reads a repository, and at the headline sheet only for requests you have checked stay at or below 100K — one of our three four-file prompt audits crossed. Among the subagent transcripts in our data that crossed 100K, the median first crossing was the 13th or 14th call. For short, repeated jobs like the prompt audit, Haiku 5.5 is both cheaper and clearly stronger than Haiku 4.5.

What Anthropic shipped on 7 October

Haiku 5.5 (claude-haiku-5-5) has a 1M-token context window, 128K output tokens, adaptive thinking on by default with the same effort levels as Opus 5.5 and Sonnet 5.5, and the newer tokenizer introduced with Opus 4.7. Its price depends on prompt length:

Rate, per million tokensPrompt up to 100,000 tokensPrompt over 100,000 tokensHaiku 4.5
Input$0.10$0.50$1.00
Output$0.50$2.50$5.00
5-minute cache write$0.125$0.625$1.25
1-hour cache write$0.20$1.00$2.00
Cache read$0.01$0.05$0.10

No other current model has a step like this. Opus 5.5, Sonnet 5.5 and the Fable models charge the same per-token rate across their 1M window. The documentation also notes that Haiku 5.5’s tokenizer produces “approximately 30% more tokens” than Haiku 4.5 for the same text, and that “the exact increase depends on the content.”

Claude Code 2.1.293, published the same day, adds the model, makes it what the haiku alias resolves to on the Anthropic API unless ANTHROPIC_DEFAULT_HAIKU_MODEL says otherwise, and lists both rates in its changelog entry. Anthropic’s developer account describes the model as pairing “well with Opus 5.5 or Sonnet 5.5 as a subagent.” The same day’s platform release notes also cut Sonnet 5.5’s cache-read price from $0.20 to $0.10 per million tokens; we use the new price below.

The Hacker News thread on the launch argued about exactly the question in this article. One commenter called 100K “an absurdly low cutoff” for agents; another replied that the complaint was vibes, since Anthropic’s own figure is 50% cheaper above the line. Both are right about the arithmetic. What was missing was a measurement of where real traffic falls.

Where the 100K line falls on real traffic

Method. Every Claude Code transcript under ~/.claude/projects/ on one macOS machine, written between 10 July and 7 October 2026: 26 main-session transcripts and 182 subagent transcripts that contain at least one API call, after dropping 32 synthetic error records that carry no tokens. Each assistant turn records a usage object with uncached input, cache-write, cache-read and output token counts. We took one record per message.id — Claude Code writes the same assistant message more than once, and summing every record double-counts, which we learned the hard way in September — and defined prompt size as uncached input plus cache writes plus cache reads for that call. That is the quantity Anthropic’s rate table calls the prompt; the cache-read and cache-write rows in the table only make sense if cached tokens count toward it, and the prompt-audit run below shows a call crossing on cache reads alone.

All but 16 of these calls ran on Opus 5, Opus 5.5, Opus 4.8, Fable 5 or Fable 5.1, which use the same newer tokenizer as Haiku 5.5, so the token counts transfer. The sessions are real work: writing, research, operations and the test harnesses behind this site, most of them long-running and autonomous. One machine is not a sample of anything wider. It is an existence proof of a shape.

TranscriptsCallsMedian promptCalls over 100KInput-side tokens over 100K
All (208)12,901209,57177.9%94.8%
Main sessions (26)8,871286,66190.2%97.9%
Subagents (182)4,030101,89150.7%74.5%

The distribution for all calls: 22.1% of prompts were 100K or under, 25.8% between 100K and 200K, 37.3% between 200K and 500K, and 14.8% over 500K. The largest was 987,697 tokens. Subagents were smaller — 49.3% at or under 100K, 39.0% between 100K and 200K, 11.6% between 200K and 500K, and one call in a thousand above that — but still half over the line.

Sessions that cross do so early. Of the 26 main-session transcripts, 18 crossed 100K at some point, and among those the median first crossing was the 10th or 11th call (quartiles: the 6th and the 24th). The 8 that never crossed were short. Of the 182 subagent transcripts, 82 crossed, at a median of the 13th or 14th call, and three quarters of them had crossed by their 19th call. The subagent is the role Anthropic suggests for this model, and in these transcripts a subagent that crossed did so well inside a run that was already longer than the median subagent’s 11 calls.

The token mix was the one we reported in September: 0.018% uncached input, 6.6% cache writes and 93.4% cache reads. That mix is why the 100K step bites so hard. On a cached prompt the dominant line is cache reads, and cache reads are $0.01 below the line and $0.05 above it.

What the step does to the discount

Pricing the same 12,901 calls on each sheet, with 1-hour cache writes charged at the 1-hour rate and 5-minute writes at theirs:

Price sheetModeled cost of the same token countsAgainst Haiku 4.5’s rates
Haiku 4.5$833.14—
Haiku 5.5, headline rate on every call$83.31−90.0%
Haiku 5.5, actual two-rate sheet$395.80−52.5%
Sonnet 5.5, with the 7 October cache price$1,339.14+61%
Opus 5.5$2,678.27+221%

Calls over 100K carry $390.61 of the $395.80, or 98.7%. The modeled Haiku 5.5 cost is 4.75 times what the headline rate implies. Split by role, main sessions would come out 51.4% below Haiku 4.5’s rates and subagents 63.5% below.

Two caveats on the baseline. These sessions ran on Opus and Fable; we priced their token counts and did not run them on Haiku, and a different model may take more or fewer turns to do the same work. And 52.1% of the calls had prompts over 200K, which is Haiku 4.5’s context limit, so the Haiku 4.5 column is what its price list gives for this traffic, not a bill that model could have produced.

Haiku 5.5 is still far cheaper than the models these sessions actually ran on — about a seventh of Opus 5.5 on the same tokens. The step does not make Haiku expensive. It makes the advertised saving over Haiku 4.5 something you only get on short prompts, and on agent traffic the saving is closer to the 50% Anthropic quotes for long prompts than to the 90% in the headline.

Then there is the tokenizer. The token counts above came from models that already use the newer tokenizer, so they are the counts Haiku 5.5 would see. Haiku 4.5 would have counted the same text as fewer tokens. Under a uniform 30% token-inflation assumption, Anthropic’s figure, the Haiku 4.5 side of the comparison drops to about $640.88 and the modeled saving on this traffic is 38.2%. At the 64.3% we measured on an unnatural filler sample, it would be about 22%. Our two natural-text samples sit between those, and neither establishes the inflation for a whole corpus.

The same text is more tokens

We sent the same document to both models through claude -p and compared the prompt-token totals, subtracting each model’s total on a one-line request (21,416 tokens on Haiku 5.5, 22,263 on Haiku 4.5, almost all of it Claude Code’s system prompt and tool definitions) to isolate the document. The wrapper differs by a few hundred tokens between runs, so treat these as approximate.

SampleHaiku 4.5 tokensHaiku 5.5 tokensChange
Three source files, 16 KB (Python and Astro)5,9257,593+28%
One documentation page, 16,151 words of Markdown prose30,59242,760+40%
100,028 words of filler built from a 30-word vocabulary121,151199,045+64%

Code came in near Anthropic’s figure, prose above it, and the filler well above — a reminder that the increase depends on content, and that repetitive text is an unfair test. Two consequences follow. A prompt that was 80K tokens on Haiku 4.5 may be over 100K on Haiku 5.5, and per-token prices alone overstate the saving on any workload where you hold the text constant.

Claude Code prices the step correctly, and haiku now means Haiku 5.5

Two checks on Claude Code 2.1.293, both with claude -p --output-format json, which reports total_cost_usd and a per-model breakdown:

The cost figure knows both rates. A one-line request on claude-haiku-5-5 produced a 21,416-token prompt (10,933 tokens written to the one-hour cache, 10,481 read) and a reported cost of $0.00229 — the headline sheet to the cent. A request carrying the filler document produced a 220,461-token prompt (207,350 cache-write, 13,109 cache-read, 2 uncached) and a reported cost of $0.20801, which is the over-100K sheet to the cent; the headline sheet would give $0.04160. Which tokens count toward the threshold shows up better in the prompt-audit run below: its seventh call had 2 uncached tokens, 1,426 written to cache and 99,664 read from cache, 101,092 in all, and Claude Code priced it at the over-100K sheet, consistent with the rate table’s separate over-100K price for cache reads. Claude Code computes these figures locally at list price, so the /usage block and the status line will be right for Haiku 5.5 unless your organization pays contracted rates, in which case the documented modelPricing setting still applies. Anthropic’s Console bill, not the local figure, is the final word.

The alias moved. claude -p --model haiku reported its usage under claude-haiku-5-5. That is the documented behavior on the Anthropic API when ANTHROPIC_DEFAULT_HAIKU_MODEL is not set, and it reaches further than /model haiku: a subagent defined with model: haiku, the built-in claude-code-guide agent, and anything routed through CLAUDE_CODE_SUBAGENT_MODEL=haiku now run on Haiku 5.5. The model configuration page adds that if your model setting is haiku, a session saved on a Haiku model resumes on whatever haiku resolves to now. Other providers lag: on Bedrock, Google Cloud and Foundry the alias still means Haiku 4.5. To stay on the old model deliberately on the Anthropic API, pin claude-haiku-4-5-20251001.

Haiku 5.5 on the prompt audit: nine for nine

On 4 October we tested /doctor prompt-audit on a small repository with nine planted problems and six correct instructions kept as controls. Opus 5.5 and Sonnet 5.5 found all nine in every run; Opus listed no controls and Sonnet listed one, at low confidence. Haiku 4.5 found five, seven and six, missed a nonexistent npm script and a missing deploy script every time, and listed the documented ultrathink keyword as a problem every time.

We reran the same fixture today, three times each on claude-haiku-5-5 and claude-haiku-4-5-20251001, both on Claude Code 2.1.293, with the same harness and scoring rules.

Planted problemHaiku 4.5, 4 OctHaiku 4.5, 7 OctHaiku 5.5, 7 Oct
1. “think step by step”3/32/33/3
2. All-caps pressure3/32/33/3
3. Nonexistent npm run test:unit0/31/33/3
4. Missing scripts/deploy.sh0/30/33/3
5. Tabs vs 2 spaces3/32/33/3
6. Missing @docs/architecture.md1/31/33/3
7. <thinking> tags2/32/33/3
8. “As a large language model…“3/33/33/3
9. Retired model ID3/33/33/3
Found per run5, 7, 63, 5, 89, 9, 9
Controls listed as findings3 (ultrathink, every run)4 (ultrathink every run; npm test once)0

Haiku 5.5’s reports read like the Opus and Sonnet ones did. Each run read package.json to show that test:unit does not exist and that AGENTS.md already says npm test, checked that scripts/ and docs/ are absent, kept ultrathink on purpose as “a thinking keyword the agent acts on,” and handled the tabs-versus-spaces contradiction as a flag for the user rather than picking a side. One run added a low-confidence note that the reviewer subagent’s one-line brief might be too thin, which is not one of our controls and which we did not count either way.

Haiku 4.5 on the same day was worse than on 4 October and more variable. Its first run found three problems. Two of its three runs proposed changing AGENTS.md to tabs even though both source files use two spaces, as it had done before. The third run flagged npm test as a command that does not exist — it does; package.json defines test — alongside ultrathink, which it again wanted replaced with an API parameter.

Model, 7 OctoberMean cost per runRangeMean timeAPI calls per runOutput tokens per run, of which thinking
Haiku 4.5$0.22$0.18–$0.2569 s11–144,071–7,432, of which 1,925–3,840
Haiku 5.5$0.04$0.03–$0.0696 s5–715,510–24,489, of which 11,088–18,717

Haiku 5.5 made fewer calls, wrote three to five times as many output tokens, most of them thinking, and took longer. It was 83% cheaper per run even though one run showed the 100K step in miniature: on its seventh call the prompt reached 101,092 tokens, and that run cost $0.057 against $0.027 and $0.033 for the other two. Priced at the headline rates throughout, it would have been $0.035; the one call over the line, which also carried 8,201 output tokens at $2.50 instead of $0.50, added $0.022.

What to do

  1. Budget Haiku 5.5 by prompt length, not by the headline. For anything that reads a repository — a main session, an Explore-style subagent, a review agent — assume the over-100K sheet: $0.50 input, $2.50 output, $0.05 cache read. Use the headline sheet only for requests you have verified stay at or below 100K. Even our four-file prompt audit crossed in one run of three.
  2. Check what haiku now points at in your setup. On the Anthropic API, model: haiku on a subagent or CLAUDE_CODE_SUBAGENT_MODEL=haiku moved those agents to Haiku 5.5 with 2.1.293, unless you override the alias. Measure a few of their transcripts the way we did above before assuming the saving.
  3. Recount tokens before comparing. The same input text cost 28% to 40% more tokens in our samples, so recount any document or context budget you carry over from Haiku 4.5. Output and thinking are a separate question: on the same audit, Haiku 5.5 wrote three to five times as many output tokens as Haiku 4.5.
  4. Trust the Claude Code cost figure for Haiku 5.5. It applies the right sheet per call. If you reconcile against Anthropic’s Console, the Console is the bill; Claude Code’s figure is list price.
  5. Use Haiku 5.5 for the prompt audit. It found all nine planted problems in every run at $0.04. Our 4 October advice to switch to Opus or Sonnet for the audit applied to Haiku 4.5 and no longer does on this test.

What we didn’t test

One machine, one kind of work. The 12,901 calls are the traffic one machine produced, most of it long autonomous sessions. A developer who starts fresh sessions every hour will have smaller prompts than ours. The subagent figures are the more portable ones, and they still come from one machine.

Almost all of the transcripts ran on Opus and Fable, not on Haiku. We priced their token counts; we did not run those sessions on Haiku 5.5, and over half of the calls were larger than Haiku 4.5’s 200K window. A different model may take more or fewer turns to do the same work. The default auto-compaction threshold on a 1M window, roughly 967K tokens, does nothing to keep requests inside the cheaper tier.

The tokenizer samples are three documents. Code, prose and filler, measured through Claude Code’s prompt wrapper rather than the token-counting endpoint. The direction is clear; the exact percentages are approximate, and they say nothing about output tokens.

Quality was measured on one small fixture. Nine planted problems in four instruction files, three runs per model. Haiku 5.5’s nine-for-nine says nothing about agentic coding at large, and Anthropic’s benchmark claims are theirs, not ours.

Costs are list prices as Claude Code reports them. Our runs used a Max subscription, where claude -p draws from the plan’s usage limits; we did not measure how either tier consumes a subscription allowance. API users see the Console bill.

Companion reading

Sources

  1. Anthropic — Introducing Claude Haiku 5.5, 7 October 2026. “On average, it now costs around 75% less to run,” and “priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens.”
  2. Anthropic — Pricing. Both Haiku 5.5 rows, Haiku 4.5, Opus 5.5 and Sonnet 5.5 rates, and “Claude Haiku 5.5 is priced by prompt length.”
  3. Anthropic — Claude Haiku 5.5 overview. The per-rate table by prompt length, 1M context, 128K output, default effort medium.
  4. Anthropic — What’s new in Claude Haiku 5.5. “The same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5. The exact increase depends on the content.” Adaptive thinking on by default.
  5. Anthropic — Claude Platform release notes, 7 October 2026. Haiku 5.5 launch; Sonnet 5.5 cache reads cut from $0.20 to $0.10 per million tokens.
  6. Anthropic — Claude Code model configuration. The alias table by provider, ANTHROPIC_DEFAULT_HAIKU_MODEL, “Use v2.1.293 or later with Haiku 5.5,” a session saved on Haiku 4.5 resuming on Haiku 5.5 when the model setting is haiku, and “A Haiku 5.5 request costs more per token when its prompt is longer than 100K tokens.”
  7. Anthropic — Claude Code changelog, 2.1.293: “Added Claude Haiku 5.5 (claude-haiku-5-5), now the default Haiku model on the Anthropic API — 1M context, $0.10/$0.50 per Mtok ($0.50/$2.50 for prompts over 100K).”
  8. Anthropic — Manage costs effectively. Claude Code computes cost figures locally from token counts at list price; the modelPricing setting for contracted rates.
  9. Anthropic — Subagents. model: haiku in subagent frontmatter and CLAUDE_CODE_SUBAGENT_MODEL.
  10. Claude Developers on X — Haiku 5.5 in Claude Code, 7 October 2026. “It pairs well with Opus 5.5 or Sonnet 5.5 as a subagent.”
  11. Hacker News — Claude Haiku 5.5, 7 October 2026; 654 points and 328 comments by 6 pm on 7 October, Pacific time.
  12. npm — @anthropic-ai/claude-code. 2.1.293 published 7 October 2026, 17:18 UTC.

Our own measurements, all on 7 October 2026 with Claude Code 2.1.293 on macOS: the transcript scan described above (208 transcripts, 12,901 distinct assistant messages after deduplication by message.id and after dropping 32 synthetic error records, 10 July to 7 October); five claude -p runs for the price and alias checks; six claude -p "/doctor prompt-audit ." runs on the 4 October fixture, three each on claude-haiku-5-5 and claude-haiku-4-5-20251001, scored by the rules in that article; and six claude -p runs for the tokenizer samples. Everything ran with a clean environment, --strict-mcp-config, and the two unrelated plugins disabled.

FAQ

Is Claude Haiku 5.5 really 90% cheaper than Haiku 4.5? Per token, yes, but only on requests whose prompt is 100,000 tokens or shorter. Above that, every rate is five times higher — half of Haiku 4.5’s price rather than a tenth. On 12,901 real Claude Code API calls from one machine, 94.8% of input-side tokens were in prompts over 100K, so the recorded token counts priced on Haiku 5.5’s sheets come to $395.80 against $833.14 at Haiku 4.5’s rates: 52.5% less, not 90%. Those are modeled costs; the sessions ran on Opus and Fable. Haiku 5.5 also counts the same text as more tokens — roughly 30% by Anthropic’s figure, 28% to 40% in our samples. Under a uniform 30% token-inflation assumption, the modeled saving on this traffic falls to 38.2%.

What counts toward Haiku 5.5’s 100,000-token threshold? The whole prompt, including tokens served from the cache. Anthropic’s rate table lists separate over-100K prices for cache reads and cache writes. In one of our prompt-audit runs, the seventh call had 2 uncached tokens, 1,426 cache-write tokens and 99,664 cache-read tokens — 101,092 in all — and Claude Code 2.1.293 priced it at the over-100K sheet. A 220,461-token request was priced the same way: $0.208, exactly what the higher rates give, against $0.042 at the headline rates. Anthropic’s Console bill, not Claude Code’s local figure, is the final word.

Does Claude Code’s cost display know about Haiku 5.5’s two price levels? Yes, from Claude Code 2.1.293. The changelog entry for Haiku 5.5 lists both rates, and in our checks the cost reported by claude -p --output-format json matched the over-100K sheet to the cent on a 220,461-token prompt and the headline sheet on a 21,416-token prompt. Claude Code computes these figures locally at list price, so a contracted rate still needs the modelPricing setting.

Does the haiku alias in Claude Code now mean Haiku 5.5? On the Anthropic API, yes, from 2.1.293, unless ANTHROPIC_DEFAULT_HAIKU_MODEL overrides it. We ran claude -p --model haiku and the usage breakdown reported claude-haiku-5-5. That covers /model haiku, subagents with model: haiku, and CLAUDE_CODE_SUBAGENT_MODEL set to haiku. Anthropic’s documentation adds that if your model setting is haiku, a session saved on Haiku 4.5 resumes on Haiku 5.5. On Bedrock, Google Cloud and Microsoft Foundry the alias still means Haiku 4.5. To keep the old model on the Anthropic API, pin the full ID claude-haiku-4-5-20251001.

Is Haiku 5.5 good enough to run Claude Code’s prompt audit? On our test repository, yes. Haiku 5.5 flagged all nine planted problems in three runs out of three and listed none of the six control instructions as a problem. Its nine of nine matched what Opus 5.5 and Sonnet 5.5 did on 4 October; its zero control findings matched Opus, while Sonnet had listed one. Haiku 4.5, rerun the same day on the same Claude Code version, found three, five and eight. Haiku 5.5 averaged $0.04 and 96 seconds per run; Haiku 4.5 averaged $0.22 and 69 seconds. One of the three Haiku 5.5 runs crossed 100K tokens on its seventh call and cost about twice the other two.

How many more tokens does Haiku 5.5’s tokenizer produce? Anthropic says approximately 30% more than Haiku 4.5 for the same text, depending on content. We measured three samples through Claude Code: three source files came out 28% larger, a 16,000-word documentation page 40% larger, and a repetitive filler document 64% larger. The extra tokens also push prompts toward the 100K line sooner.

Was this helpful?

Related reading


Reviews independently produced · Editorial policy

Read more reviews →