On September 22, 2026, Anthropic introduced Claude Opus 5.5. About 90 minutes later, OpenAI announced GPT-6 Sol and Luna. The three models arrived almost at once, but their benchmark charts leave the most useful question for product teams unanswered: With the same budget, which system gets more work done?
Opus 5.5 brings high-end capability below the price of Opus 5. OpenAI keeps Astra for its hardest work, places Sol in the middle, and prices Luna for large volumes. The contest goes beyond which model is smarter. Both companies want to become the default choice when businesses put agents to work every day.
TL;DR
- API prices per 1M input/output tokens are $2/$10 for Sol, $0.10/$0.50 for Luna, and $4/$20 for Opus 5.5. Those are unit prices, not the cost of finishing a job.
- OpenAI reports a 33.2% score for Sol on AutomationBench at $0.27 per task. Its comparison is against Opus 5, however, not the Opus 5.5 released that day.
- Artificial Analysis finds that Sol and Luna cost substantially less to run than the 5.6 generation, while their overall capability scores are largely flat. Some knowledge-work results declined.
- Opus 5.5 topped the Artificial Analysis Intelligence Index at launch. Anthropic cut prices too, so this is more than a simple matchup between a powerful model and a cheap one.
- The practical way to choose is to measure accepted results, correction costs, and completion time on the workflows users actually run.
One launch day, two ways to win the market
OpenAI cut Sol's price from $4/$20 to $2/$10 per 1M input/output tokens. Luna fell from $0.20/$1.20 to $0.10/$0.50. For products that call a model constantly, a small saving on each call can change the cost of millions of calls. Luna's output price fell by more than half; it would be inaccurate to say every part of its pricing fell by exactly 50%.

Anthropic also lowered its price. Opus 5.5 costs $4/$20, or 20% less per token than Opus 5. Anthropic says the cost per task at default settings fell by about 40% because the model uses fewer tokens. That is Anthropic's reported result, not a guaranteed saving on every workflow. Still, both launches make clear where the competition is heading: the operating budget for agents.

The pricing and product positioning suggest two different approaches. Anthropic wants customers to hand difficult tasks to Opus with less rework. OpenAI covers more budget levels: Astra for the most demanding tasks, Sol for agents that run often, and Luna for simpler tasks in large numbers. This is an interpretation of their positioning, not inside information about either company's strategy.
Why price per million tokens is only part of the story
An agent may need to read documents, call tools, check its output, and retry when something fails. If the result still cannot be used, a person must fix it. Token spend is therefore only one part of the bill. Suppose model A costs $0.10 per run but succeeds on half its tasks; model B costs $0.16 per run and succeeds on 90%.
If attempts are independent and retries are allowed, their expected costs per successful task are $0.20 and about $0.18, respectively. These are hypothetical figures, not Sol or Opus results. They show how the cheaper model on a pricing page can cost more in production.
In AutomationBench results published by OpenAI, Sol at xhigh scored 33.2% at $0.27 per task. Opus 5 at max scored 26.9% and cost about 11.1 times as much as Sol. That is a strong signal for Sol's efficiency on this multi-step workflow test. But the opponent is Opus 5, not the newly released Opus 5.5. This chart cannot tell us how Sol and Opus 5.5 would perform on the same job.

What independent testing confirms, and what it does not
Artificial Analysis tested Sol and Luna with its own evaluation suite. Average cost per task on its Intelligence Index was $1.06 for Sol max, versus $1.99 for GPT-5.6 Sol max. Luna max cost $0.07, versus $0.18 for its predecessor. Both new models actually used slightly more output tokens; lower prices drove most of the savings. These are costs on a benchmark suite, not fixed prices for customer tasks.
Then there is quality. Artificial Analysis found that Intelligence Index scores were broadly flat against the 5.6 generation. Sol max gained two points on the Coding Agent Index to reach 57, while Luna max lost two points and landed at 41. Both models scored lower on GDPval-AA v2.1, a knowledge-work evaluation. In AA-Briefcase, evaluators found some deliverables with weaker presentation or missing requested items. If an agent writes client reports, a missing section may take more time to repair than the saved tokens are worth.

The reduction in wrong answers needs the same care. On AA-Omniscience, Sol max's hallucination rate fell from 92% to 60%, but the share of questions it attempted fell from 99% to 83%. Reported accuracy also fell from 59% to 54%. In research, admitting uncertainty is often better than making something up. The system still needs a way to find more evidence or pass the question to a human. A refusal alone does not finish the job.
Is Opus 5.5 moving into the midrange itself?
Anthropic priced Opus 5.5 input and output at twice Sol's rates. Yet a model call is not a finished task. If Opus gets a job right in one pass while Sol needs three attempts and manual edits, Opus may be cheaper per usable result. If both handle a simple job on the first try, Sol has a clear price advantage. Both outcomes are plausible, which is why the specific workload matters.
Anthropic reports strong Opus 5.5 performance on coding and knowledge-work tasks. Artificial Analysis also ranked it first on the Intelligence Index at launch. In Anthropic's own chart, Opus 5.5 scored 40.0% on AutomationBench, but that chart does not include the new Sol. Putting Anthropic's 40.0% beside OpenAI's 33.2% as though they were a direct head-to-head test would be misleading. Different settings, tools, and runs can change the result.
Luna makes scale possible, but errors still have a cost
Luna raises a different question: What happens when a small task must be repeated hundreds of thousands of times? At $0.10/$0.50 per 1M input/output tokens, it is attractive to test Luna as the first screening layer. OpenAI also reports that Luna max reached 66.6% on DeepSWE v1.1, at a per-task cost 93% or 96% lower than selected Opus 5 and Fable 5 configurations. Those competitors are older models, not Opus 5.5. Artificial Analysis got different DeepSWE results in its independent runs, so figures from the two evaluations should not be combined into a single ranking.
One clear use is to have Luna filter data, assign labels, or compile a shortlist of news worth reading. Sol can then compare multiple sources, while ambiguous cases or costly mistakes go to Opus and a human reviewer. In a crypto news monitoring system, for example, the first layer could remove duplicate listing announcements; interpreting token unlock terms would require a separate verification step. This is a proposed design, not the result of testing all three models in one system. If a screening error discards important news, Luna's low price cannot save the workflow.
A model-selection playbook for product and research teams
Start with 30–50 real tasks: easy cases, ambiguous ones, and examples the existing system has already mishandled. Define acceptance criteria first. Are the sources right? Is every required section present? Does the code pass tests? Run each model on the same cases with comparable tools and data. Record retries, manual editing time, waiting time, and the final cost of each accepted result.
Only then route the work: try Luna on tasks that are easy to check, Sol on recurring multi-step work, and Opus on difficult exceptions. Run a one-week pilot with human review before scaling up. When models or prices change, rerun the same cases.
For another example of testing a model's promise against its real use, read Whales' analysis of Jev. For AI projects with tokens, setting a testable standard before putting money at risk also resembles Whales' trading-thesis playbook: write down what would prove the thesis wrong.
What could break the price-led market-share thesis?
This thesis weakens if benchmark results fail to translate into usable work. A model may perform well on a test yet repeatedly get a ticker wrong, miss a contract condition, or produce code that needs extensive repairs. In that case, correcting its mistakes consumes the savings on tokens.
Prices will change too. Anthropic has already cut Opus 5.5's price and said Sonnet and Haiku 5.5 will follow. The September 22 price chart is a snapshot. If a competitor cuts prices further or gets better at difficult tasks, the best routing strategy will change. AI tokens raise another question: Will users saving money on inference increase a project's revenue or the value captured by its token? There is no automatic link.
Conclusion
Opus 5.5 brings high-end capability below the price of Opus 5. Sol and Luna expand the range of jobs that agents can handle at scale. Independent data confirms lower costs for Sol and Luna on published evaluations, but it also shows that quality has not improved across the board. That is why the contest should be judged by completed work, not pricing or benchmark scores alone.
Opus 5.5 is worth testing on difficult cases that take time to fix. Sol is worth testing as the default for multi-step work. Luna is worth testing for screening when there is a clear checking step. After a week on real data, look at one calculation: total money and human effort divided by accepted results. The model that wins on that measure is the one that truly earns the budget.
FAQ
Is GPT-6 Sol more capable than Claude Opus 5.5?
There is no basis for a blanket claim. Some OpenAI charts compare Sol against older Claude models, while Artificial Analysis ranks Opus 5.5 highly on overall capability. Compare them on the same tasks, tools, effort settings, and acceptance criteria.
Can GPT-6 Luna replace Sol for every agent?
That should not be the default assumption. Luna costs much less, but some independent coding and knowledge-work scores declined. It is better suited to tasks where errors are easy to spot and hard cases can be passed to a stronger model.
Does a 50% API price cut mean a 50% lower bill?
Not necessarily. The bill depends on token usage, caching, the number of calls, and the human effort needed to correct results. Measure the cost of each correctly completed task.
Why does OpenAI say Sol beats Opus while Anthropic says Opus 5.5 leads?
The companies use different comparisons. OpenAI pits Sol against Opus 5 in some charts, while Opus 5.5 arrived later. Test both new models on the same tasks, tools, and scoring criteria.
Which model should a team test first if it can pick only one?
For an agent that runs often, start by testing Sol. For hard work that takes time to repair, test Opus 5.5. For large volumes of easy-to-check work, try Luna and escalate hard cases to a stronger model. The final choice should depend on the cost per correctly completed task.
Do lower AI costs automatically benefit AI project tokens?
No. Lower costs may allow a project to serve more users, but its token benefits only if that activity creates value through the token's actual mechanism. Examine revenue, repeat usage, and how value flows back to the token instead of inferring token performance from a model pricing chart.