Back
DeepSeekAugust 01, 20262 sources

DeepSeek releases V4-Flash-0731, intensifying AI's 'race to zero'

AI Analysis

DeepSeek officially released V4-Flash-0731 on July 31, a mixture-of-experts model with 284B total parameters and 13B active, that scored 82.7% on Terminal-Bench — surpassing its own V4-Pro-Preview — and hit 50 on the ArtificialAnalysis Index, a single point behind GLM-5.2 and GPT-5.6 Luna. Priced at roughly $0.14 per million input tokens and $0.28 per million output, it is aggressively positioned to set a new economic baseline for enterprise-grade agentic automation. Axios and Nikkei framed the release, alongside a retrained V4 lineup, as accelerating AI's 'race to zero,' with V4-Pro output tokens reportedly near $0.87/M versus $50/M for Anthropic's Fable.

The efficiency story is central to the community reaction. On r/LocalLLaMA, the Hugging Face release post (773 upvotes) and a 'going to cause another market crash' thread (600+ upvotes) drove intense excitement, while r/DeepSeek asked 'how is that even possible?' (356 upvotes) about a ~300B model undercutting far smaller ones on cost. Hugging Face's Victor Mustar deployed a free public inference endpoint within hours.

Crucially for adoption, V4-Flash speaks both OpenAI's Responses API and Anthropic's Claude-compatible message format, letting developers swap it into existing agent stacks with minimal rewiring. That interoperability, plus a claimed 1-million-token context, positions it as a drop-in cost-cutter rather than a walled garden.

The release also arrived amid a broader Chinese-lab surge — MiniMax's open-weight H3 video model and ByteDance's Seedance 2.5 landed the same week — signaling coordinated pricing and capability pressure on U.S. incumbents. The open question is durability: whether DeepSeek's economics hold at scale under real production load, and how OpenAI and Anthropic respond beyond their own recent price cuts.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog