Back to Blog
Tools & Resources7 min readSeptember 28, 2026

OpenAI and Anthropic Both Cut Flagship Prices on 22 September

GPT-6 Sol and Luna launched at half the price of GPT-5.6, and Claude Opus 5.5 arrived 20% cheaper than Opus 5 on the same day. What the new rate cards mean for your margins.

Sarah Chen

Sarah Chen

Content at NeedBase

On 22 September 2026, the two largest model vendors cut prices on the same day. OpenAI released GPT-6 Sol at $2 input / $10 output per million tokens and GPT-6 Luna at $0.10 / $0.50. Anthropic released Claude Opus 5.5 at $4 / $20, down from $5 / $25 for Opus 5. If you have an AI feature in production, your cost per request probably just changed, and so did the maths on which model you should be routing to.

The new rate cards

GPT-6 Sol โ€” $2 / $10, versus $4 / $20 for GPT-5.6 Sol. That is exactly half on both input and output. OpenAI pitches it for code review, debugging, data analysis and complex recurring work.

GPT-6 Luna โ€” $0.10 / $0.50, versus $0.20 / $1.20 for GPT-5.6 Luna. Half on input, 58.3% cheaper on output. Positioned for summarisation, extraction and straightforward questions.

Claude Opus 5.5 โ€” $4 / $20, versus $5 / $25 for Opus 5, a 20% cut. The bigger change is caching: cache reads fall from $0.50 to $0.20 per million tokens, a 60% cut, and cache writes from $6.25 to $5. Anthropic says typical workloads come out about 40% cheaper than on Opus 5 once caching is counted, and that output is around 30% faster.

OpenAI told VentureBeat the GPT-6 prices are permanent, not introductory, and attributed the cuts to caching and inference improvements. OpenAI also offers a 90% discount on cached input reads. GPT-6 Astra, the top tier, keeps its own higher pricing.

How the field lines up now

Using the figures VentureBeat published alongside the launch, the mid-tier is suddenly crowded. GPT-6 Sol and Claude Sonnet 5 are now both $2 / $10. Grok 4.7 is $2 / $6 for prompts up to 200K tokens. Gemini 3.8 Flash is $0.75 / $3.75 until 31 December 2026, then $1.50 / $7.50 โ€” worth noting if you are choosing it for its price today. At the very bottom, GPT-6 Luna at $0.10 / $0.50 sits alongside Xiaomi's MiMo-V2.6-Flash at $0.14 / $0.28.

And Opus 5.5 now costs exactly what GPT-5.6 Sol cost last week. Anthropic claims it beats its own larger Fable model on many benchmarks; TechCrunch reports the company calls it the strongest model it has tested. Treat vendor benchmark claims as a reason to test, not a reason to switch.

The details that will bite

Opus 5.5 has breaking changes. Anthropic's announcement says thinking can no longer be disabled on Opus 5.5, and a "preserved thinking" safeguard applies to API accounts created after 31 August 2026. If your integration turns thinking off to save tokens or latency, swapping the model string is not enough โ€” test before you migrate, and budget for the thinking tokens you were previously avoiding.

Cheaper per token is not cheaper per task. A model that thinks longer or writes more verbose output can cost more per completed job even at a lower rate. OpenAI's own figure for Sol on AutomationBench is $0.27 per task, which is the kind of number you should be measuring on your own workload.

Caching is now the biggest lever. Both vendors pushed the discount on cached reads. If your prompts start with a long, stable system prompt, tool definitions or a knowledge block, and you are not structuring them to be cached, you are leaving most of this price cut on the table.

What to actually do this week

Re-run your cost model with the new numbers. Take last month's token counts per feature and price them against the new cards. For many products the answer will be a real margin improvement with no code change beyond a model string.

Run a small eval before switching. Pick 50 to 100 real requests from your logs, run them through the current model and the candidate, and compare quality and cost per completed task. An afternoon of this beats a week of reading benchmark tables.

Revisit your routing. If you route easy requests to a cheap model and hard ones to a flagship, both ends just moved. Luna may now handle work you were sending to a mid-tier model, and Opus 5.5 may now be affordable for work you were keeping off Opus.

Decide whether to pass savings on. If you price per seat, this is margin. If you price per usage or credits, competitors will reprice, and customers will notice.

Keep your provider abstraction honest. Same-day cuts from two vendors are a reminder that the cheapest option changes monthly. If switching models takes more than a config change and an eval run, fix that first.

The bottom line

On 22 September, GPT-6 Sol and Luna halved OpenAI's mid and low-tier prices permanently, and Claude Opus 5.5 took 20% off Anthropic's flagship with a 60% cut to cache reads. Re-price your AI features against the new cards, run a short eval on your own traffic, and check Opus 5.5's thinking changes before you flip the switch.

Found this useful?

Share it with a founder who needs it.

Ready to launch your product?

Join thousands of makers who launched on NeedBase.

Submit Your Product โ†’