ModelsArchitectures & capability
Claude Haiku 5.5 matches GPT-6 Luna's list price, but independent tests find it costs about three times more per task at max effort
Anthropic's Claude Haiku 5.5 is listed at $0.10/$0.50 per million tokens, the same as GPT-6 Luna. Independent measurements, however, show heavy reasoning-token use at max effort and a tokenizer that counts more tokens, so per-task cost exceeds what list prices suggest.

Haiku 5.5 and Luna have the same list price, but at max effort Artificial Analysis measured Haiku 5.5 using about 162K output tokens per task, roughly 129K of them for reasoning, against about 50K for Luna. That works out to about $0.21 per task for Haiku 5.5 versus $0.07 for Luna, which cuts against Anthropic's claim that the model is about 75% cheaper than Haiku 4.5. [4] [7]
Some costs don't show up in the per-token price. Haiku 5.5 uses the newer tokenizer shared with Sonnet and Opus 5.5, which counted about 1.25x more tokens than Haiku 4.5 for the same text. Prompts over 100K tokens are billed at $0.50 input and $2.50 output, five times the base rate. Luna's long-prompt surcharge starts later, at 272K tokens, and is smaller. [1] [3]
The effort setting changes cost and latency a lot. In one SVG test, low effort took 7 seconds and cost under a tenth of a cent, while max effort took more than five minutes and cost about 3.4 cents. Going from xhigh to max added 2 points on the Intelligence Index for roughly 1.8x the tokens. At max effort, the first answer token took about 323 seconds to arrive. [1] [4] [7]
The large gains over Haiku 4.5 are vendor-reported. Those include OSWorld 2.1 rising from 15.7% to 72.4% and Terminal-Bench 4.0 from 0% to 39.2%. The Terminal-Bench figure is the max-effort result; at medium effort the score is about 20%. Artificial Analysis independently gave the max-effort setting an Intelligence Index score of 43. [3] [7] [8]
independent testing at max effort found $0.21 per task for Haiku 5.5 versus $0.07 for Luna. Implication: comparing models on per-token price can understate real spend.
Read the full assessment
Measuring cost per completed task, and testing lower effort levels first, gives a more accurate budget.
Executive brief
Independent testing complicates Anthropic's "75% cheaper" claim. Claude Haiku 5.5 has the same list price as OpenAI's GPT-6 Luna, $0.10/$0.50 per million tokens. But at maximum effort, Artificial Analysis measured it using about 162K output tokens per task against Luna's 50K. That puts its average cost at $0.21 per task versus $0.07 for Luna (The Decoder). Anthropic's own benchmark gains over Haiku 4.5 are large, but they are vendor-reported. Simon Willison also found that the new tokenizer counts about 1.25x more tokens for the same text, which works as a hidden price rise (Willison).
What changed and event timeline
Haiku 4.5 sets the baseline
It was priced at $1/$5 per million tokens and had no reasoning-effort levels. Willison later called its pelican SVG poor.
Opus 5.5 and GPT-6 Sol/Luna launch about an hour apart
Luna was priced at $0.10/$0.50, half the price of GPT-5.6 Luna. Opus 5.5 came in at $4/$20.
Sonnet 5.5 ships
Anthropic claimed it runs 30%+ faster and costs up to 30% less for most work, at unchanged list price. It became the free-tier model on claude.ai.
Haiku 5.5 released
It arrived alongside a 50% cut to Sonnet 5.5 cache-read prices and new monthly API credits for Max and Team subscribers.
08: Independent token-use measurements published
Artificial Analysis gave the max-effort setting an Intelligence Index score of 43, but flagged it as very verbose.
Capabilities and access
Read the full section
- Model ID:
claude-haiku-5-5. Available on the Claude Platform, AWS, Google Cloud and Azure (Anthropic). - Price per million tokens, prompts up to 100K tokens: $0.10 input, $0.50 output, $0.01 cache reads.
- Price above 100K tokens: $0.50 input, $2.50 output (Anthropic).
- Context window: 1M tokens, up from 200K (The Decoder).
- Reasoning effort: low, medium, high, xhigh and max. Reasoning cannot be turned off, and the default is medium (Willison).
Technical analysis for researchers and developers
- Tokenizer: Haiku 5.5 uses the newer tokenizer shared with Sonnet and Opus 5.5 (Anthropic). Per-token price comparisons therefore understate real cost.
- Cost and latency swing widely with effort level. Max effort took 5 minutes 9 seconds and cost 3.38¢.
- Max effort is mostly reasoning tokens. About 129K of the 162K output tokens per task go to reasoning.
Read the full section
- Tokenizer: Haiku 5.5 uses the newer tokenizer shared with Sonnet and Opus 5.5 (Anthropic). Willison's token counter found about 1.25x more tokens than Haiku 4.5 for the same prompt. Per-token price comparisons therefore understate real cost.
- Cost and latency swing widely with effort level. In Willison's pelican test, low effort took 7 seconds and cost 0.0936¢. Max effort took 5 minutes 9 seconds and cost 3.38¢.
- Max effort is mostly reasoning tokens. About 129K of the 162K output tokens per task go to reasoning. Moving from xhigh to max adds 2 index points for roughly 1.8x the tokens (The Decoder).
- Throughput versus first-token latency. Artificial Analysis measured 243 tokens/sec but about 323 seconds to the first answer token at max effort (Artificial Analysis).
- Tooling: llm-anthropic 0.30 can now pick up new models with
llm anthropic refresh, without a plugin release.
Claims and evidence
- OSWorld 2.1: 72.4% (Haiku 4.5: 15.7%) — Vendor-reported ()
- Terminal-Bench 4.0: 39.2% (Haiku 4.5: 0.0%)
- GDPval-AA: 1,620 vs Luna 1,437; OSWorld: 72.4% vs Luna 48.9% — Vendor-reported ()
Read the full section
| Claim | Status |
| OSWorld 2.1: 72.4% (Haiku 4.5: 15.7%) | Vendor-reported (Anthropic) |
| Terminal-Bench 4.0: 39.2% (Haiku 4.5: 0.0%) | Vendor-reported; this is the max-effort score, and medium effort scores about 20% (VentureBeat) |
| GDPval-AA: 1,620 vs Luna 1,437; OSWorld: 72.4% vs Luna 48.9% | Vendor-reported (VentureBeat) |
| Humanity's Last Exam: 45.9% without tools, 57.4% with tools | Vendor-reported |
| About 75% cheaper than Haiku 4.5 on average | Vendor-reported. Independent per-task costs point the other way when compared with Luna |
| Intelligence Index 43; $0.21 per task | Independent (Artificial Analysis) |
| Hallucination rate 40% vs Luna 77% | Reported by The Decoder, citing Artificial Analysis |
Context and prior work
Read the full section
- Haiku 4.5 cost 10x Luna's price, and Willison describes it as "showing its age" (Willison).
- The release follows two weeks of mutual price cuts across the Opus, Sonnet and GPT-6 lines (Willison).
- This is the first Haiku-class model with adjustable effort (XDA).
- Willison flags a possible problem with his long-running pelican benchmark: Haiku 5.5's max-effort reasoning trace explicitly recognized it as "the classic pelican-on-bicycle SVG test."
Limitations, safety and contested findings
- Safety claims are vendor-only. No independent corroboration was found (Anthropic).
- The cost picture is contested. Luna reaches similar performance at about one-third the cost per task (The Decoder).
- Long prompts cost far more. Above 100K tokens, input price rises 5x. Luna's surcharge starts later, at 272K, and is smaller: $0.20/$0.75 (Willison).
Read the full section
- Safety claims are vendor-only. Anthropic reports "major improvements" on alignment evaluations and stricter cyber safeguards than Haiku 4.5, though it says more defensive tasks are allowed than on Sonnet 5.5. No independent corroboration was found (Anthropic).
- The cost picture is contested. Luna reaches similar performance at about one-third the cost per task (The Decoder).
- Long prompts cost far more. Above 100K tokens, input price rises 5x. Luna's surcharge starts later, at 272K, and is smaller: $0.20/$0.75 (Willison).
- Max effort can backfire. Willison reported that Sonnet and Opus 5.5 can exhaust their token budget at max effort without producing usable output (Willison).
Business and practitioner implications
- Budget per completed task, not per token, and test low and medium effort before defaulting higher.
- Keep prompts under 100K tokens where possible. For long-context workloads, Luna is cheaper on list price (Willison).
- Subscriber credits: Max 5x gets $100/month, Max 20x gets $200, and Team gets up to $500 pooled.
Read the full section
- Benchmark your own traffic. Budget per completed task, not per token, and test low and medium effort before defaulting higher.
- Keep prompts under 100K tokens where possible. For long-context workloads, Luna is cheaper on list price (Willison).
- Subscriber credits: Max 5x gets $100/month, Max 20x gets $200, and Team gets up to $500 pooled. Credits do not roll over, and turning off auto-reload stops requests when the balance runs out (Claude support).
- Early customer reports are vendor-curated. Anthropic's launch quotes cite Box at half the latency and Asana at over 30% lower task-completion latency (VentureBeat).
The source trail.
Sources (9)
Simon Willison: Claude Haiku 5.5
simonwillison.netPelicans and a bunch of notes on the pricing of the new Claude Haiku 5.5 https://t.co/F3Sob8WvsG
Related coverage; assess separately
x.com