Back
DeepSeekAugust 01, 20261 sources

DeepSeek V4-Flash Ships at Record $0.14/M Input Tokens, Up to 100x Cheaper Than Claude Fable 5

AI Analysis

DeepSeek V4-Flash-0731 is the sharpest expression yet of China's cost-compression strategy. At $0.14/M input and $0.28/M output tokens, users on r/LocalLLaMA calculated it at roughly 105x lower total cost than Claude Fable 5 on identical tasks. The model carries 284 billion parameters, runs in about 142GB of GPU memory at FP4, and — notably — beats DeepSeek's own 1.6-trillion-parameter V4-Pro by 14% on intelligence leaderboards, a striking demonstration that architecture and training data can outweigh raw scale.

On agentic benchmarks it posts strong numbers, surpassing V4-Pro-Preview on Terminal Bench 2.1 with a score of 82.7, and adds native Responses API and Codex support. The API entered public beta on July 31, and llama.cpp quickly added MTP/DSpark support, cementing local-inference enthusiasm.

The competitive impact is immediate and loud. An r/Anthropic thread with over 1,400 upvotes argued V4-Flash is 18x cheaper on input and 28x on output while 'matching Opus 4.8,' demanding Anthropic at least cut Sonnet pricing. Community members described a psychological shift: 'token anxiety is gone — developers now budget based on task cost, not token fear.' Coupled with Alibaba's Qwen 3.8-Max the same week, V4-Flash intensifies the narrative that Chinese labs are dictating the price floor for frontier-adjacent capability.

Caveats: leaderboard parity with Opus is a community claim, not an independently audited result, and 'matches Opus' on aggregate scores can mask task-specific gaps — a separate r/Anthropic thread titled 'As a scientist Fable is not usable' shows real workloads still diverge. What to watch: whether Western labs respond with price cuts (AWS already slashed GPT-5.6 pricing) and whether V4-Flash's quality holds under sustained agentic use.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog