Learn
Claude Haiku 5.5 vs GPT-6 Luna: Same Price, Different Fit
Claude Haiku 5.5 and GPT-6 Luna both list at $0.10/$0.50 per million tokens, but Haiku 5.5 charges five times more above 100,000-token prompts and leads in Anthropic's agent benchmarks. Compare prices, benchmarks, platforms and best-fit jobs.
Separate adjacent ideas before you evaluate them. Use this page when similar names or layers sound interchangeable but lead to different decisions.
Editorial guide
Guide
Start with the core separation before you compare workflows, pricing, or plans.
Anthropic released Claude Haiku 5.5 on October 7, 2026, and for prompts up to 100,000 tokens it lists at exactly the price of OpenAI's GPT-6 Luna: $0.10 per million input tokens and $0.50 per million output tokens, with $0.01 cache reads and $0.125 cache writes on both. The sticker is the same, so the choice turns on three other things: what each model gets right per task, what happens to the price on long prompts, and the platform and API rules around each one. This guide uses only the vendors' published figures.
The short answer
- Choose Claude Haiku 5.5 for computer-use and browser agents, for knowledge-work and visual tasks where Anthropic's published results favor it, for subagents that run under Claude Sonnet 5.5 or Opus 5.5, and when your stack already runs on the Claude API, Amazon Bedrock, Google Cloud or Microsoft Foundry.
- Choose GPT-6 Luna when individual prompts regularly exceed 100,000 tokens, when you need OpenAI's Responses API tools, Flex processing or EU data residency, or when your agents already run on GPT-6 Sol in the OpenAI stack.
- For short classification, routing and extraction, the list price is identical, so decide on accuracy measured on a labeled sample of your own data.
Price and limits side by side
Item | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
Input / output per 1M tokens | $0.10 / $0.50 for prompts up to 100,000 tokens | $0.10 / $0.50 for prompts up to 272,000 input tokens |
Long prompts | $0.50 / $2.50 for prompts over 100,000 tokens | 2x input and cache rates and 1.5x output above 272,000 input tokens ($0.20 / $0.75) |
Cache read / 5-minute cache write | $0.01 / $0.125 ($0.05 / $0.625 over 100,000 tokens) | $0.01 / $0.125 |
Batch discount | 50% | 50% on Batch and Flex |
Context window / max output | 1M / 128K tokens | 1.05M / 128K tokens |
Reasoning control | Adaptive thinking with effort levels, | |
Reliable knowledge cutoff | June 2026 | May 18, 2026 |
API model ID | | |
Anthropic's long-prompt price applies to the whole request once its prompt passes 100,000 tokens, and OpenAI's multipliers apply to the whole request once it passes 272,000 input tokens. Both price steps are per request, not per extra token.
What long prompts do to the bill
The examples below use list prices without caching. They are illustrations, not measurements.
Single request | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
20,000 input + 1,000 output tokens | $0.0025 | $0.0025 |
150,000 input + 2,000 output tokens | $0.080 | $0.016 |
400,000 input + 5,000 output tokens | $0.213 | $0.084 |
Below 100,000 tokens the two cost the same per token. Between 100,000 and 272,000 tokens, Haiku 5.5 charges five times its base rate while Luna stays at its base rate, so a 150,000-token document costs five times as much to read on Haiku 5.5. Above 272,000 tokens Luna's rate rises too, but less, and Haiku 5.5 still costs about 2.5 times as much in the example. If you want Haiku 5.5 for long documents, split them into chunks under 100,000 tokens.
Tokenizers also differ. Anthropic says Haiku 5.5 splits text into roughly 30% more tokens than Haiku 4.5 did, and Luna uses OpenAI's separate tokenizer, so one document can produce different token totals on the two models. Run a sample of real inputs through both APIs and compare the reported usage before committing a budget.
The benchmarks Anthropic published
Anthropic's launch table compares Haiku 5.5 directly with GPT-6 Luna. These are Anthropic's own runs, described in the Haiku 5.5 system card, and the OSWorld figures use its offline subset. Sonnet 5.5 is included only for reference.
Benchmark | Claude Haiku 5.5 | GPT-6 Luna | Claude Sonnet 5.5 (reference) |
|---|---|---|---|
GDPval-AA v2.1, knowledge work | 1620 | 1437 | 1840 |
AA-Briefcase v1.1, knowledge work | 1578 | 1336 | 1824 |
OSWorld 2.1 offline subset, computer use | 72.4% | 48.9% | 83.9% |
Terminal-Bench 4.0, agentic coding | 39.2% | 16.4% | 70.6% |
FrontierCode 1.1 Main, agentic coding | 46.4% | 42.4% | 52.1% (Xhigh) |
Chartography without tools, visual reasoning | 46.4% | 29.1% | 61.6% |
Haiku 5.5 is ahead on each of the six rows it shares with Luna in this table. The gaps are widest on computer use and Terminal-Bench 4.0 and narrowest on FrontierCode, the mergeable-code test. OpenAI has not published Luna results on these benchmark versions; its own launch post reports Luna at 66.6% on DeepSWE v1.1 at max effort, a test Anthropic's table does not include. Treat the table as a reason to shortlist Haiku 5.5 for agent-style work, not as a verdict on your workload. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding.
Which model fits which job
Workload | Better starting point | Why |
|---|---|---|
Short classification, routing and tagging | Either; test both | Same list price; accuracy on your labels decides |
Computer use and browser agents | Claude Haiku 5.5 | Larger lead in Anthropic's OSWorld results; browser use tool on the Claude API and Google Cloud |
Documents over 100,000 tokens per request | GPT-6 Luna | No price step until 272,000 tokens |
Coding subagents | Depends on the lead model | Haiku 5.5 pairs with Sonnet 5.5 or Opus 5.5; Luna pairs with GPT-6 Sol |
Cost-sensitive batch jobs | Either | Both take 50% off on Batch; Luna also offers Flex at 50% |
Platform and API differences
Claude Haiku 5.5 runs on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. It is also generally available in GitHub Copilot on Pro, Pro+, Max, Business and Enterprise, billed at Anthropic's list price. In Claude Code, the haiku alias selects it on the Anthropic API from version 2.1.293. Moving from Haiku 4.5 needs code changes: non-default temperature, top_p or top_k values, assistant prefill and manual budget_tokens all return errors, and responses can start with thinking blocks.
GPT-6 Luna runs in the OpenAI API through the Responses, Chat Completions and Batch endpoints, with web search, file search, computer use and MCP tools on the Responses API. EU data residency is available on Standard, Flex and Batch processing, and regional processing adds 10%. Chat Completions supports function calling only with reasoning effort set to none, and Luna cannot be fine-tuned.
Bottom line
At the same $0.10/$0.50 list price, GPT-6 Luna is the cheaper model whenever single prompts pass 100,000 tokens, and Claude Haiku 5.5 is the stronger candidate in Anthropic's published agent and knowledge-work results. For short, high-volume work, run a labeled sample through both and keep the one with more accepted outputs per dollar on your own data.
Evidence boundary
Official sources
Editorial guidance grounded in official product sources.
FAQ
Common questions
Is Claude Haiku 5.5 cheaper than GPT-6 Luna?
Not for short prompts: both list at $0.10 input and $0.50 output per million tokens, with $0.01 cache reads. Haiku 5.5 rises to $0.50/$2.50 for prompts over 100,000 tokens, while Luna keeps its base rate until 272,000 input tokens and then applies 2x input and 1.5x output. The tokenizers differ, so compare bills on a sample of your own text.
Which model posts better benchmark results?
In Anthropic's launch table Haiku 5.5 leads Luna on all six shared rows, for example 72.4% against 48.9% on OSWorld 2.1 offline and 39.2% against 16.4% on Terminal-Bench 4.0. These are Anthropic's own runs; OpenAI has not published Luna results on these benchmark versions.
When does Haiku 5.5's long-prompt price apply?
When a request's prompt is longer than 100,000 tokens. That request is billed at $0.50 per million input tokens, $2.50 per million output tokens, $0.05 for cache reads and $0.625 for five-minute cache writes. Keeping chunks under 100,000 tokens keeps the lower rates.
Can both models drive computer-use agents?
Yes. GPT-6 Luna supports the computer use tool in the Responses API. Haiku 5.5 supports computer use through the computer_toolset_20260801 toolset on the Claude API and Google Cloud, plus a browser use tool there, and it posts the higher OSWorld 2.1 result in Anthropic's table.
Where can I run each model?
Haiku 5.5 runs on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and in GitHub Copilot on paid plans. GPT-6 Luna runs in the OpenAI API as gpt-6-luna, with EU data residency on Standard, Flex and Batch processing.
Do the two models accept the same request parameters?
No. Haiku 5.5 returns an error for non-default temperature, top_p or top_k values, for assistant prefill and for manual budget_tokens, and it is steered with the effort parameter. Luna is steered with reasoning.effort from none to max, and on Chat Completions it supports function calling only with reasoning effort set to none.
Next steps
Open both sides of the distinction
Open the most relevant product pages or follow-up guides for each side of the distinction after the split is clear.