DeepSeek to raise API prices after processing 8 trillion tokens in a single day

DeepSeek told users to 'plan accordingly' ahead of a 'significant' increase in API prices across all services — a striking reversal for the company that ignited the global AI price war. The trigger was overwhelming demand for V4-Flash, released barely a week earlier and described by one research firm as by far the cheapest of the well-known models to run, at $0.14 per million input tokens and $0.28 per million output. DeepSeek says it processed 8 trillion tokens on August 1 alone, exposing the economics of sustaining aggressive pricing at scale.
Agent builders read the hike as a broader signal that inference compute is hitting saturation and cost pressure is arriving from multiple directions at once — ending DeepSeek's role as price-war disruptor. On r/DeepSeek, threads titled 'price increase panic' and 'party over?' captured the frustration, with some users pivoting the conversation toward local open-source alternatives, while a countervailing 'Stop whining, Start Supporting' thread urged patience.
The move coincides with DeepSeek resuming an $8 billion funding round at a roughly $74 billion valuation, suggesting the company is trading rock-bottom pricing for sustainable margins and a war chest. It also fits a wider Chinese cost-and-monetization recalibration: Alibaba is testing revenue-sharing for its open-weight Qwen line the same week.
For developers who built products on DeepSeek's near-zero inference costs, the hike is a live risk to unit economics. Watch the exact new price schedule, whether V4-Flash remains cheaper than Western rivals after the increase, and whether the pricing reset drives migration to Qwen, Kimi K3 or self-hosted open models.