Oct 8 edition/Reporting & analysis
ModelsBusinessAgents

ModelsArchitectures & capability

Claude Haiku 5.5 matches GPT-6 Luna's list price, but independent tests find it costs about three times more per task at max effort

Anthropic's Claude Haiku 5.5 is listed at $0.10/$0.50 per million tokens, the same as GPT-6 Luna. Independent measurements, however, show heavy reasoning-token use at max effort and a tokenizer that counts more tokens, so per-task cost exceeds what list prices suggest.

Illustration from Simon Willison’s Weblog: Claude Haiku 5.5 matches GPT-6 Luna's list price, but independent tests find it costs about three times more per task at max effort
Image: Simon Willison’s Weblog — Original article ↗
THE CORE IDEAS4 TAKEAWAYS
01

Haiku 5.5 and Luna have the same list price, but at max effort Artificial Analysis measured Haiku 5.5 using about 162K output tokens per task, roughly 129K of them for reasoning, against about 50K for Luna. That works out to about $0.21 per task for Haiku 5.5 versus $0.07 for Luna, which cuts against Anthropic's claim that the model is about 75% cheaper than Haiku 4.5. [4] [7]

02

Some costs don't show up in the per-token price. Haiku 5.5 uses the newer tokenizer shared with Sonnet and Opus 5.5, which counted about 1.25x more tokens than Haiku 4.5 for the same text. Prompts over 100K tokens are billed at $0.50 input and $2.50 output, five times the base rate. Luna's long-prompt surcharge starts later, at 272K tokens, and is smaller. [1] [3]

03

The effort setting changes cost and latency a lot. In one SVG test, low effort took 7 seconds and cost under a tenth of a cent, while max effort took more than five minutes and cost about 3.4 cents. Going from xhigh to max added 2 points on the Intelligence Index for roughly 1.8x the tokens. At max effort, the first answer token took about 323 seconds to arrive. [1] [4] [7]

04

The large gains over Haiku 4.5 are vendor-reported. Those include OSWorld 2.1 rising from 15.7% to 72.4% and Terminal-Bench 4.0 from 0% to 39.2%. The Terminal-Bench figure is the max-effort result; at medium effort the score is about 20%. Artificial Analysis independently gave the max-effort setting an Intelligence Index score of 43. [3] [7] [8]

WHY IT MATTERS

independent testing at max effort found $0.21 per task for Haiku 5.5 versus $0.07 for Luna. Implication: comparing models on per-token price can understate real spend.

Read the full assessment

Measuring cost per completed task, and testing lower effort levels first, gives a more accurate budget.

Executive brief

Independent testing complicates Anthropic's "75% cheaper" claim. Claude Haiku 5.5 has the same list price as OpenAI's GPT-6 Luna, $0.10/$0.50 per million tokens. But at maximum effort, Artificial Analysis measured it using about 162K output tokens per task against Luna's 50K. That puts its average cost at $0.21 per task versus $0.07 for Luna (The Decoder). Anthropic's own benchmark gains over Haiku 4.5 are large, but they are vendor-reported. Simon Willison also found that the new tokenizer counts about 1.25x more tokens for the same text, which works as a hidden price rise (Willison).

What changed and event timeline

  1. Haiku 4.5 sets the baseline

    It was priced at $1/$5 per million tokens and had no reasoning-effort levels. Willison later called its pelican SVG poor.

  2. Opus 5.5 and GPT-6 Sol/Luna launch about an hour apart

    Luna was priced at $0.10/$0.50, half the price of GPT-5.6 Luna. Opus 5.5 came in at $4/$20.

  3. Sonnet 5.5 ships

    Anthropic claimed it runs 30%+ faster and costs up to 30% less for most work, at unchanged list price. It became the free-tier model on claude.ai.

  4. Haiku 5.5 released

    It arrived alongside a 50% cut to Sonnet 5.5 cache-read prices and new monthly API credits for Max and Team subscribers.

  5. 08: Independent token-use measurements published

    Artificial Analysis gave the max-effort setting an Intelligence Index score of 43, but flagged it as very verbose.

Capabilities and access

  • Model ID: claude-haiku-5-5. Available on the Claude Platform, AWS, Google Cloud and Azure (Anthropic).
  • Price per million tokens, prompts up to 100K tokens: $0.10 input, $0.50 output, $0.01 cache reads.
  • Price above 100K tokens: $0.50 input, $2.50 output (Anthropic).
Read the full section
  • Model ID: claude-haiku-5-5. Available on the Claude Platform, AWS, Google Cloud and Azure (Anthropic).
  • Price per million tokens, prompts up to 100K tokens: $0.10 input, $0.50 output, $0.01 cache reads.
  • Price above 100K tokens: $0.50 input, $2.50 output (Anthropic).
  • Context window: 1M tokens, up from 200K (The Decoder).
  • Reasoning effort: low, medium, high, xhigh and max. Reasoning cannot be turned off, and the default is medium (Willison).

Technical analysis for researchers and developers

  • Tokenizer: Haiku 5.5 uses the newer tokenizer shared with Sonnet and Opus 5.5 (Anthropic). Per-token price comparisons therefore understate real cost.
  • Cost and latency swing widely with effort level. Max effort took 5 minutes 9 seconds and cost 3.38¢.
  • Max effort is mostly reasoning tokens. About 129K of the 162K output tokens per task go to reasoning.
Read the full section
  • Tokenizer: Haiku 5.5 uses the newer tokenizer shared with Sonnet and Opus 5.5 (Anthropic). Willison's token counter found about 1.25x more tokens than Haiku 4.5 for the same prompt. Per-token price comparisons therefore understate real cost.
  • Cost and latency swing widely with effort level. In Willison's pelican test, low effort took 7 seconds and cost 0.0936¢. Max effort took 5 minutes 9 seconds and cost 3.38¢.
  • Max effort is mostly reasoning tokens. About 129K of the 162K output tokens per task go to reasoning. Moving from xhigh to max adds 2 index points for roughly 1.8x the tokens (The Decoder).
  • Throughput versus first-token latency. Artificial Analysis measured 243 tokens/sec but about 323 seconds to the first answer token at max effort (Artificial Analysis).
  • Tooling: llm-anthropic 0.30 can now pick up new models with llm anthropic refresh, without a plugin release.

Claims and evidence

  • OSWorld 2.1: 72.4% (Haiku 4.5: 15.7%) — Vendor-reported ()
  • Terminal-Bench 4.0: 39.2% (Haiku 4.5: 0.0%)
  • GDPval-AA: 1,620 vs Luna 1,437; OSWorld: 72.4% vs Luna 48.9% — Vendor-reported ()
Read the full section
ClaimStatus
OSWorld 2.1: 72.4% (Haiku 4.5: 15.7%)Vendor-reported (Anthropic)
Terminal-Bench 4.0: 39.2% (Haiku 4.5: 0.0%)Vendor-reported; this is the max-effort score, and medium effort scores about 20% (VentureBeat)
GDPval-AA: 1,620 vs Luna 1,437; OSWorld: 72.4% vs Luna 48.9%Vendor-reported (VentureBeat)
Humanity's Last Exam: 45.9% without tools, 57.4% with toolsVendor-reported
About 75% cheaper than Haiku 4.5 on averageVendor-reported. Independent per-task costs point the other way when compared with Luna
Intelligence Index 43; $0.21 per taskIndependent (Artificial Analysis)
Hallucination rate 40% vs Luna 77%Reported by The Decoder, citing Artificial Analysis

Context and prior work

  • Haiku 4.5 cost 10x Luna's price, and Willison describes it as "showing its age" (Willison).
  • The release follows two weeks of mutual price cuts across the Opus, Sonnet and GPT-6 lines (Willison).
  • This is the first Haiku-class model with adjustable effort (XDA).
Read the full section
  • Haiku 4.5 cost 10x Luna's price, and Willison describes it as "showing its age" (Willison).
  • The release follows two weeks of mutual price cuts across the Opus, Sonnet and GPT-6 lines (Willison).
  • This is the first Haiku-class model with adjustable effort (XDA).
  • Willison flags a possible problem with his long-running pelican benchmark: Haiku 5.5's max-effort reasoning trace explicitly recognized it as "the classic pelican-on-bicycle SVG test."

Limitations, safety and contested findings

  • Safety claims are vendor-only. No independent corroboration was found (Anthropic).
  • The cost picture is contested. Luna reaches similar performance at about one-third the cost per task (The Decoder).
  • Long prompts cost far more. Above 100K tokens, input price rises 5x. Luna's surcharge starts later, at 272K, and is smaller: $0.20/$0.75 (Willison).
Read the full section
  • Safety claims are vendor-only. Anthropic reports "major improvements" on alignment evaluations and stricter cyber safeguards than Haiku 4.5, though it says more defensive tasks are allowed than on Sonnet 5.5. No independent corroboration was found (Anthropic).
  • The cost picture is contested. Luna reaches similar performance at about one-third the cost per task (The Decoder).
  • Long prompts cost far more. Above 100K tokens, input price rises 5x. Luna's surcharge starts later, at 272K, and is smaller: $0.20/$0.75 (Willison).
  • Max effort can backfire. Willison reported that Sonnet and Opus 5.5 can exhaust their token budget at max effort without producing usable output (Willison).

Business and practitioner implications

  • Budget per completed task, not per token, and test low and medium effort before defaulting higher.
  • Keep prompts under 100K tokens where possible. For long-context workloads, Luna is cheaper on list price (Willison).
  • Subscriber credits: Max 5x gets $100/month, Max 20x gets $200, and Team gets up to $500 pooled.
Read the full section
  • Benchmark your own traffic. Budget per completed task, not per token, and test low and medium effort before defaulting higher.
  • Keep prompts under 100K tokens where possible. For long-context workloads, Luna is cheaper on list price (Willison).
  • Subscriber credits: Max 5x gets $100/month, Max 20x gets $200, and Team gets up to $500 pooled. Credits do not roll over, and turning off auto-reload stops requests when the balance runs out (Claude support).
  • Early customer reports are vendor-curated. Anthropic's launch quotes cite Box at half the latency and Asana at over 30% lower task-completion latency (VentureBeat).
FOLLOW THE EVIDENCE

The source trail.

Sources (9)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief