DeepSeek launches V4.1 Flash multimodal beta at 333-400+ tok/s

DeepSeek's V4.1 Flash beta is the new fact this cycle — distinct from the V4-Pro 0813 reasoning model and its 3.6x price hike covered previously. V4.1 Flash adds native multimodal support, decode speeds of 333–400+ tokens per second, and a 1-million-token context, all at pricing comparable to the earlier V4-Flash, with a 60% cache-hit price cut taking effect September 10. That speed-and-context combination positions it as a high-throughput workhorse for agentic and long-document workloads rather than a frontier reasoning showcase.
Testers were quick to caveat: the model ID reportedly carries an 'expires-on-0910' tag, which several flagged as a short-window A/B test rather than a final release, and no independent benchmark reproduction had appeared. That tempers the raw-speed enthusiasm on r/LocalLLaMA, where the tok/s numbers drew praise.
The broader story is DeepSeek's self-sufficiency push against US export controls. Bloomberg reported the company plans at least 160,000 Huawei Ascend 950DT accelerators at a gigawatt-scale Inner Mongolia data center for inference — a bet on domestic Chinese silicon that, if it works, would insulate DeepSeek from NVIDIA supply constraints. The 150 senior backend engineering roles it opened signal it's scaling infrastructure to match. Competitively, V4.1 Flash keeps DeepSeek in the open-weight throughput race against Qwen, Kimi and Meta's Muse Spark, even as its pricing strategy has grown less predictable (the V4-Pro hike drew backlash from cost-sensitive builders who adopted it precisely for cheap inference). The durable question is whether Huawei Ascend hardware can deliver the price-performance that made DeepSeek's early models disruptive; the answer determines whether its self-sufficiency bet is a strength or a ceiling.