The price of artificial intelligence stopped moving in one direction this week. On Thursday, the Chinese lab DeepSeek released the full version of its flagship V4-Pro model and sharply raised what it charges developers to use it, with some rates climbing by as much as 1,100 per cent. Days before that, OpenAI and Google had cut theirs. The AI price war, which for a year and a half meant one thing only, that tokens kept getting cheaper, has split into two streams running in opposite directions.
The direction of travel matters because a whole industry has been built on the assumption that it only goes one way. Thousands of startups, many of them in India, wrote their cost models around the idea that the price of a million tokens would keep falling. DeepSeek was the reason it fell so fast. When the company arrived in early 2025 with models priced at a fraction of what American labs charged, it forced everyone below its own line. Now the cheapest name in the market is the one raising prices.
What DeepSeek changed, and then changed back
DeepSeek posted a notice on its developer platform on 6 August telling customers to “plan usage accordingly” ahead of what it called a significant increase. The scale became clear this week. From 16 August, according to the company’s pricing page, output tokens on its V4-Flash model will cost 1.32 dollars per million at peak hours, up from a flat 28 cents. Its more capable V4-Pro model rises to 3.96 dollars at peak, from 87 cents. Input prices climb too, and the company has added a peak and off-peak structure, charging less during quieter hours to push demand away from its busiest windows.
DeepSeek said it was adjusting prices “to allocate resources more reasonably,” language that points at the real constraint. The bottleneck is not profit margin. It is capacity. One tracker recorded DeepSeek’s Flash model processing eight trillion tokens in a single day earlier this month, a volume spike that landed on hardware already running hot. Raising the price is the oldest way there is to ration a scarce thing.
The company insists it remains the cheapest option even after the rise. Its founder, Jun Song, said on the platform X that even after an increase of between two and ten times, DeepSeek would still undercut most Western rivals. The arithmetic supports him. Even at the new peak rate, DeepSeek’s flagship sits far below Anthropic’s Fable 5, which a report by Fortune put at 50 dollars per million output tokens.
The Western labs are cutting, hard
While DeepSeek climbed, the American labs went the other way. OpenAI cut the price of its GPT-5.6 Luna model by 80 per cent on 30 July and its mid-tier Terra model by 20 per cent, moves widely read as pressure on the Chinese challengers. A report in Caixin Global described DeepSeek’s price rise as “a break from the prolonged price war” among China’s AI developers, coming just as OpenAI leaned the other way. Google, too, shipped a cheaper workhorse model into the same market overnight.
So the same week produced a Chinese lab retreating from rock-bottom pricing and American labs racing toward it. That is not a contradiction. It is two companies with different problems. DeepSeek is capacity-bound and rationing. OpenAI and Google are capacity-rich, or spending enough to look it, and using price to take share.
A market splitting by tier, not just by flag
The split is not only Chinese against American. It runs through the product ladder inside each lab. The cheapest tiers are getting cheaper and being pushed at high-volume, everyday work. The frontier tiers are holding their price and adding paid speed on top. Between them, the middle is where the competition is fiercest, because that is where production workloads actually live. OpenAI’s Terra cut and DeepSeek’s V4-Pro rise are aimed at the same band of customers from opposite sides.
The real cost is moving off the model
Behind the token price, a bigger shift is under way, and it is where the money now sits. The cost of using AI is migrating away from the model itself and toward the infrastructure that runs it: the speed, the chips, the capacity, the power.
OpenAI made the point on 13 August. It previewed an Ultrafast tier that runs its flagship Sol model on hardware built by Cerebras, reaching up to 750 output tokens a second and up to 14 times its standard speed. The intelligence of the model does not change. Only the speed does. Access is limited to a small group of customers while capacity expands, and OpenAI has published no price for it. The company has already committed billions to Cerebras for low-latency compute, and Cerebras reported its cloud revenue nearly quadrupling in the second quarter as it locked in hundreds of megawatts of capacity.
That is the tell. When a lab charges nothing extra for a smarter model but reserves a premium for a faster one, the scarce resource has moved. It is no longer the intelligence. It is the silicon and the electricity underneath it. Chip makers are reading the same signal, with SMIC in China raising prices as its fabrication plants run near capacity.
Why the buildout dwarfs the price cuts
The sums involved make the token price look like a rounding error. Purchase and lease commitments by the largest American technology companies, Alphabet, Microsoft, Amazon, Nvidia, Oracle and Meta, run to roughly 1.5 trillion dollars, according to a Financial Times analysis. That is the number the price war is really about. A lab can afford to give away cheaper tokens if it believes the returns will come from owning the compute layer that everyone else has to rent.
For the buyer, the lesson of the week is that a token price is not a promise. DeepSeek’s own experience proves it. A permanent discount it introduced in May lasted only until demand forced a rethink. Any business that built its plan around one provider’s cheapest tier now knows how quickly that floor can move. The practical response is the one that always follows pricing volatility: route work between models by task, keep cheap models for simple jobs and reserve the expensive ones for hard reasoning, and never wire a whole product to a single vendor’s floor.
What it means for India
India has more riding on this than most. The country’s AI startups and its large services firms have leaned heavily on cheap inference to make their products viable, and DeepSeek’s models sit inside a great many of them. A price rise at the bottom of the market lands directly on those margins, and the peak and off-peak structure will reward whoever reschedules the heaviest work into quieter hours. The deeper signal is the one to watch. As cost moves from the model to the compute beneath it, access to chips and power becomes the thing that decides who can build at scale. That is an argument the country is already having, over data centres, electricity and where the hardware will come from, and this week made it sharper.


