DeepSeek warns of API price hikes after V4-Flash processes 8 trillion tokens in a day

DeepSeek notified API customers of significant price increases across all its services, a sharp reversal from the aggressive price-cutting strategy that made it a symbol of low-cost Chinese AI. The trigger: its newly-launched V4-Flash model processed 8 trillion tokens in a single day, overwhelming GPU capacity and forcing DeepSeek to raise prices to manage demand rather than keep undercutting rivals.
V4-Flash itself is a 284-billion-parameter, text-only model with a 1M-token context window that entered public beta at highly competitive rates — $0.14 per 1M input tokens and $0.28 per 1M output, with a 98% discount on cached input — and achieved an intelligence index of 50, ranking second among 162 models in its class. The price hike lands almost immediately after those record-low rates were touted, which fueled community skepticism about the sustainability of ultra-cheap Chinese model economics (an r/DeepSeek thread on 'the truth behind DeepSeek's price increase' hit 468 upvotes).
Beyond pricing, DeepSeek resumed an $8 billion funding round at a roughly $74 billion valuation and took a 2.31% stake in humanoid robot maker Unitree ahead of its IPO — signaling ambitions beyond pure model APIs into robotics and embodied AI.
Skeptics read the price hike as evidence that the low-cost model was subsidized and unsustainable at scale, undercutting the narrative that Chinese labs had structurally cheaper economics. What to watch: how much prices actually rise, whether demand holds at higher rates, and whether the Unitree stake signals a broader hardware pivot.