DeepSeek launches V4-Pro-0813 with agent capabilities, then imposes 4-14x price hike

DeepSeek released DeepSeek-V4-Pro-0813 via API, app, and web with substantially improved agent and software-engineering capabilities, a 1M-token context window, native support for the OpenAI Responses API format, and low/high/max reasoning-effort tiers. Early benchmarks put it at 87.9 on Terminal Bench 2.1 and 62.7 on DeepSWE — one Reddit thread noted the 87.9 lands within a tenth of a point of Claude Fable 5's 88.0 at roughly 57x cheaper output pricing. The company also stood up a 'DeepSeek Harness Team' and shipped DeepSeek Harness v0.1 in developer preview under an MIT license, explicitly aiming to compete with Claude Code.
The headline shift, however, is economic. Effective 16:00 UTC on August 16, peak-hour V4-Pro output pricing rises from $0.87 to $3.96 per million tokens, with cache pricing reportedly up as much as 10x and overall increases community members describe as 50-1000% across the V4 family. Off-peak usage gets a 50% discount, nudging batch workloads to cheaper windows.
Competitively, this is a notable reversal for the company that arguably ignited the low-cost inference race. Coverage from Reuters and Bloomberg framed it as compute scarcity finally showing up in list prices; SCMP reported the model underwhelmed on general benchmarks but shone in cybersecurity, suggesting DeepSeek is leaning into niche specialization rather than broad frontier gains.
Skeptics on r/DeepSeek were split — some called the end of 'the free samples era,' others pointed out users aren't forced to pay DeepSeek's first-party inference prices when cheaper third-party providers host the open weights. Watch whether rivals follow with their own increases, which would confirm the scarcity narrative rather than a one-off.