Back
DeepSeekSeptember 27, 20262 sources

DeepSeek V4.1-Flash Cuts Output Price 70%, Claims to Beat Opus 5

AI Analysis

DeepSeek is pressing its cost advantage hard with V4.1-Flash, the smallest model in a new asymmetric architecture family. During off-peak hours the model prices cached input tokens at $0.003 per million and output at $0.60 per million—undercutting its own flagship V4-Pro by roughly 70% on output—while claiming to beat Anthropic's Opus 5 on select coding and reasoning benchmarks.

Technically, V4.1-Flash is a 552-billion-parameter mixture-of-experts model (released to the API September 10) built on a new asymmetric architecture with native multimodal visual understanding and a context window of up to one million tokens. The off-peak pricing wrinkle is itself a strategic signal: DeepSeek is effectively arbitraging its own idle capacity to make high-context, high-volume workloads dramatically cheaper than Western frontier APIs.

The move sharpens the week's dominant sub-theme—an accelerating LLM price war. With Meta quietly shipping Muse Spark 1.3 near $0.10 per million blended and Alibaba pushing open-weight Qwen models, the cost floor keeps dropping, and DeepSeek's 'beats Opus 5' framing is calculated to pressure premium-priced Western labs on cost-per-token.

Skeptics urge caution on the benchmark claims, which are largely self-reported, and note that off-peak pricing complicates real-world TCO comparisons. There's also a geopolitical overhang: r/singularity (218 upvotes) noted 'Dario's latest model has stirred hard feelings among the Chinese user base,' hinting at a tit-for-tat competitive mood. For developers, though, the practical takeaway is stark—frontier-adjacent capability at a fraction of incumbent prices is now routine, and Western labs increasingly compete on trust, tooling, and integration rather than raw price.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog