If you last modelled your unit economics in the spring, they are wrong by roughly half.
Average inference pricing sat at $2.04 per million tokens at the end of May. By late July it was $1.45. In early August it reached $1.16 to $1.18 โ a 43% decline over ten weeks, and the lowest figure recorded this year. ARK Invest's tracking shows an even sharper move, from $2.07 to $1.02. By BenchLM's index, frontier-level capability now costs around 12% of its March 2023 price.
A caveat before you build on those numbers: these are blended averages across models and providers, and your own bill depends entirely on which models you call and how much context you send. Check your actual invoice rather than the index. But the direction is not ambiguous and the magnitude is large.
What caused it
A price war, mostly. OpenAI cut pricing on its lower tier by 80% in late July, taking input tokens to around $0.20 per million and output to $1.20. Chinese open-weight models applied pressure from underneath, with DeepSeek reported to have cut as much as 99% earlier in the year while making weights downloadable โ which sets a ceiling on what anyone can charge for comparable capability.
Hardware efficiency is contributing, but competition is doing most of the work. That matters for how you plan, because competitive pricing can reverse and efficiency gains generally do not.
Three things this changes today
Features you costed out are now viable. Everyone building on models has a list of things rejected because the token cost per user made them unprofitable. Summarising every document on upload, enriching every record, running a quality check on every generation. Go back to that list โ a good portion of it is now cheap enough to ship, and it is the fastest source of product improvement available to you this quarter.
Your prices are probably wrong in one direction or the other. If you set a price with a margin buffer sized for spring token costs, you are either sitting on far more margin than you planned โ fine, but know it โ or you are still declining customers whose usage would now be profitable. Either way the decision was made against numbers that no longer exist.
Generous free tiers are affordable again. The reason most AI products have stingy free tiers is that free users cost real money. At current prices, a free tier that would have been reckless in May may now be your cheapest acquisition channel. Model it before dismissing it.
The trap
Cheap inference is available to your competitors on identical terms. If the only reason your product works is that tokens got cheap, you have not built an advantage, you have benefited from a subsidy that everyone received.
This cuts particularly hard for thin wrappers. When the underlying capability costs 12% of what it did three years ago and keeps falling, charging a healthy multiple for a prompt and a text box gets harder every quarter, because the gap between your price and the obvious alternative widens in full view of the customer.
What does not commoditise: proprietary data, workflow integration, distribution, and being the system of record. Those were the durable positions when inference was expensive and they are more clearly the durable positions now that it is not.
If you sell outcomes, recheck your floor
Anyone charging per result rather than per seat has a variable cost sitting under a fixed price, and that cost just moved substantially in your favour. Two consequences worth acting on.
Your worst-case customer is less bad than it was. The heavy-usage accounts that were marginal or loss-making at spring prices may now be fine, which means minimum commitments you set defensively can probably be relaxed to win deals.
And customers know prices fell. Enterprise buyers read the same coverage, and a renewal conversation in which your price is unchanged while your input costs halved is a conversation you should walk into prepared for. Better to lead with an improvement they can see than to defend a number.
What to actually do this week
Pull last month's provider invoice and divide by active users to get your real cost per user. Compare it with whatever figure your pricing was built on. Then pick the single most valuable feature you previously rejected on cost grounds and re-run the maths on it.
That is an hour of work, and for most AI products it will surface either a margin you did not know you had or a feature you can now afford to ship.
The bottom line
Token costs fell by nearly half in ten weeks and by an order of magnitude over three years. The businesses that benefit are the ones that notice and re-plan, not the ones that quietly keep the extra margin until a competitor prices against it. Recheck your costs, ship the feature you shelved, and make sure the thing customers pay you for is not simply access to a model that gets cheaper every month.