AI companies sell their models by the token, roughly three quarters of a word. You are billed for the words you send and, at a much higher rate, the words you get back.
Say your finance team asks what a million words of output costs. Two weeks ago that was a lookup. Today the honest answer needs five follow-up questions: which model, through which reseller, at what hour, on which day, and is that price still inside its promotional window.
Four things happened in six days that made this permanent.
August 16: DeepSeek Put a Clock on It
At 16:00 UTC, DeepSeek replaced the flat rate it had run since May. The current table, from their own documentation:
| deepseek-v4-pro, per million tokens | off-peak | peak |
|---|---|---|
| Text you send, repeated | $0.022 | $0.044 |
| Text you send, new | $0.66 | $1.32 |
| Text it sends back | $1.98 | $3.96 |
Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Everything else, weekends included, is off-peak at half price.
Convert that into Beijing time and peak is 09:00 to 12:00 and 14:00 to 18:00 there: the Chinese working day with a two-hour lunch cut out of the middle. That is not a pricing scheme so much as a photograph of when their own datacentre is busy.
The consequences are geographic and nobody is discussing them. North America is entirely off-peak: 01:00 UTC is 6pm the previous day in California, so every hour a US team is at their desk, DeepSeek charges the cheap rate. Europe pays for its mornings. Asia-Pacific pays for its afternoons: from here in Brisbane, peak is 11am to 2pm and 4pm to 8pm.
The model marketed on being cheap is now cheapest for the customers furthest from it.
DeepSeek framed this as off-peak being 50% lower than peak, which is true and beside the point. Reporting puts the previous flat rate for output at $0.87 per million. The new off-peak rate is $1.98. The cheapest hour under the new scheme costs more than twice the old price at any hour, and peak is over four times it.
August 17: A Discount at Two Addresses Only
The next day OpenAI started a 50% promotion on GPT-5.6 Sol, running to September 18, applied automatically. Not on its own API. On OpenRouter and Vercel AI Gateway, specifically.
— SemiAnalysis, on the Sol gateway promotionOpenRouter and Vercel account for a negligible share of OpenAI’s total token usage.
They are, however, the main public data source for estimating which models the market uses. OpenRouter publishes rankings, analysts quote them, journalists quote the analysts. Halving the price at exactly those two venues doubles volume where the scoreboard is kept and nowhere else that matters to revenue.
This is a guess about motive and SemiAnalysis presents it as one. A company might rationally discount at a reseller to win the developers who default to it. But the effect is not in dispute: the only public record of who is winning now has a deliberate surge of cheap demand built into one column.
It is not the only case. In the thread on Stripe’s acquisition, a developer noticed the same pattern from another vendor:
— zurfer, Hacker News, 19 August 2026Woah now also a 75perc discount on OpenRouter for flash 3.7. Is it really the same product (speed? and up time)? Why would Google do that?
Two days after that, Stripe agreed to buy OpenRouter.
August 21: Three Months, and Not for Subscribers
Then OpenAI cut Sol’s list price: input $5 to $4, output $30 to $20. That is 20% and 33%, putting Sol at $4/$20 against Claude Opus 5 at $5/$25, so the previously expensive option now undercuts on both.
Two conditions. It expires, listed through at least November 21. And it excludes subscribers: the cut covers pay-as-you-go, Codex credits and eligible Work plans, while Pro, Plus and Business get nothing.
That second point is the one to sit with. The cheapest way to buy Sol is now to be a developer with a credit card, and the most expensive is to be on a monthly plan. Every vendor has spent a year pushing people onto subscriptions, and the first serious price cut of the season routes around them.
Why This Works
After OpenAI cut its Luna model by 80%, TD Cowen data showed usage rose roughly 14-fold and revenue rose 34%.
Cut the price by four fifths, people use it fourteen times as much, and you collect more than before. Cheaper costs the seller nothing, because nobody is close to using as much of this as they want, and agents will consume whatever budget you give them. So a price cut is a growth tactic that looks generous, and the honest version of “we are passing on efficiency gains” is “we worked out that cheaper makes us more money.”
Which also explains DeepSeek going the other way the same week. If demand expands that easily, your problem is not price, it is running out of machines, and the fix for that is a rush-hour surcharge that pushes whatever can wait into the night.
What It Costs You
- You cannot plan on a sticker price. Any budget built on a published rate needs a date it was true and a date it stops being true. Mine had neither.
- When you run the job matters again. Bulk processing, testing and reprocessing can all wait. On DeepSeek, moving them past 10:00 UTC halves the bill. Anyone who ran overnight jobs on a mainframe would recognise this.
- Choosing the cheapest model is a live calculation at this minute, at this reseller, under this promotion. That is roughly what Stripe’s $8 billion bought.
- The “who is winning” scoreboard is contaminated. One vendor is currently paying half of one column’s price.
What This Isn’t
Promotional pricing is not sinister: electricity has had peak pricing for a century, and the only novelty is that AI was flat-rate until this month. And Sol being cheaper than Opus 5 compares stickers only. A model that takes twice as many attempts at 80% of the rate is not cheaper.
The Shape of It
Six days: a surcharge, a discount at two addresses, an acquisition of one of those addresses, and an expiring list-price cut. None of the four was a response to the others, and together they finish off the idea that a token has a price the way a litre of milk has a price.
What it has instead is a rate card with footnotes, and the footnotes are where the money is.



