DeepSeek Launches V4.1-Flash With 552B Params, Native Vision, and 60% Cheaper Output

DeepSeek released V4.1-Flash, a mixture-of-experts model with 552 billion total parameters (8B active on input, 16B on output), a 1-million-token context window, and native multimodal visual understanding. API pricing undercuts DeepSeek's own V4-Pro by up to 60%, at $0.30 per million input tokens and $1.20 per million output during peak, halved off-peak. A new architecture delivers a fourfold reduction in KV-cache memory costs — roughly 890 bytes per token, enabling 1M context on about 1GB — which infra builders on Hacker News singled out as the standout feature for long-running agents.
Community reception was loud and divided. An r/DeepSeek thread declaring 'DeepSeek v4.1 Flash is truly amazing' (184 upvotes) sat next to another calling it 'a downgrade on creativity, intelligence and state management' (141 upvotes) — the classic split between cost/efficiency gains and raw-capability regressions. The local-hosting crowd debated whether it truly beats Kimi K3.
Within hours of the MIT-licensed release to Hugging Face on September 10, developers stripped its safety layers and uploaded uncensored 'abliterated' forks that drew steady download traffic — reviving the open-weight safety-stripping debate that recurs with every capable open release. A Qwen-style heretic/uncensored GGUF variant trend was visible on r/LocalLLaMA.
Competitively, V4.1-Flash sharpens DeepSeek's cost-leader positioning against Qwen, Kimi, and Western open-weight models, and lands as DeepSeek hired CITIC Securities and a first CFO ahead of a potential Shanghai STAR Market IPO at a ~$74 billion valuation. DeepSeek was also named in Anthropic's threat report as a Claude-distilling lab, so its efficiency gains carry an IP cloud. Watch whether the KV-cache efficiency translates to real agent cost savings and how the abliterated forks factor into the open-weight-safety regulatory conversation.