DeepSeek releases V4-Flash-0731, beats its own V4-Pro on agent benchmarks at $0.14/M

DeepSeek officially shipped V4-Flash-0731, the newest member of its open-weights V4 family, retaining a 284B-total / 13B-active Mixture-of-Experts architecture but delivering large agentic gains through post-training rather than a bigger base model. It scored 82.7% on Terminal-Bench 2.1 (up from 61.8% in the preview), beat the larger V4-Pro-Preview's 72.1%, and reached parity with Gemini 3.6 Flash. Pricing is aggressive: $0.14 per million input tokens and $0.28 per million output, with a 1M-token context window and Codex compatibility. The V4 family also includes V4-Pro for advanced reasoning.
The headline mechanism is economics. Developers cited cost efficiency roughly 105x-137x lower than Claude, cache-hit reads at $0.0028/M, and described 'genuinely disruptive economics' for agentic products — about 3x cheaper than V4-Pro and 21x cheaper than Claude Sonnet 5. The weights are downloadable and runnable locally, and Hugging Face's Victor Mustar deployed a free public inference endpoint within hours of release.
Competitively, this is the sharpest move yet in the inference price war and directly pressured incumbents: OpenAI cut prices on its low- and mid-tier models the same week, and Reddit's r/Anthropic ran a 1,406-upvote thread demanding Anthropic cut Sonnet pricing. Local-LLM enthusiasts on r/LocalLLaMA (1,378 upvotes) marveled that a 300B model runs cheaper than a 9B one, and that locally-runnable models now match the top frontier model from March 2026.
Skeptical takes: developers questioned real-world capacity and whether the benchmark gains hold on large production codebases rather than toy tasks — r/DeepSeek threads explicitly asked for 'honest feedback, not hype.' What to watch: whether DeepSeek can serve demand at these prices, and whether Western labs respond with structural price cuts rather than one-off discounts.