DeepSeek's Open-Weight V4-Flash Lands at ~$0.03/Task, 105x Cheaper Than Claude Fable 5

DeepSeek officially launched V4-Flash on July 31, 2026, and by August 3 an independent research firm had crowned it the cheapest well-known AI model to run anywhere, averaging about three cents per benchmark test versus $3.15 for Anthropic's Claude Fable 5 — a roughly 105x gap. DeepSeek's published API rate is $0.14 per million input tokens and $0.28 per million output tokens, undercutting Moonshot AI's Kimi K3, OpenAI's GPT-5.6 Sol, and Anthropic's Fable 5 by wide margins.
What makes V4-Flash notable beyond price is that it is not a toy: it scores 50 on Artificial Analysis' Intelligence Index, tying Google's Gemini 3.6 Flash, and ships with enhanced agentic capabilities and open weights. That combination — frontier-adjacent intelligence, agent tooling, and near-zero marginal cost — is what has Western labs rattled about cost-per-intelligence economics.
The launch escalates China's intensifying AI price war, arriving the same week Alibaba unveiled Qwen3.8-Max. Semafor framed it as China flooding the market with cut-price, open models while U.S. labs still charge premium API rates. Notably, DeepSeek also paused its Series B fundraising round, a signal it may be prioritizing market-share capture over near-term capital.
The darker footnote: V4-Flash was quickly abused. A hacker wired the model into the open-source Hermes Agent framework to autonomously attack 460 systems via Telegram, achieving 14 intrusions — though all exploited already-patched public vulnerabilities. Community reaction on r/DeepSeek was euphoric on cost ('subscription services are dead to me,' 507 upvotes), while security researchers noted the operational failure modes mattered more than any AI-agency panic. The open question for readers: whether cheap open weights structurally erode the moat of premium closed labs, or whether tooling, evals, and harness quality remain the real differentiator.