DeepSeek-V4.1-Flash ships a causal encoder-decoder design to cut agent cache costs

DeepSeek's latest Flash model takes an architectural approach to the main cost driver in agent workloads: repeatedly re-reading large contexts. V4.1-Flash uses a causal encoder-decoder design. It activates 8B parameters per token in prefill, when the input is processed, and 16B in decode, when output is generated. That asymmetry reduces cache reads for input-heavy agentic tasks, where tool outputs and repository context far exceed generated text.
The specs are notable for an open model. V4.1-Flash has a 1,048,576-token context window and an MIT license, and it scored 74.2% on DeepSWE v1.1. That score sits between xAI's Grok 4.7 at 71.0% and Google's Gemini 4 Argon at 77.9% on the same benchmark. DeepLearning.AI's The Batch said Flash 'leapfrogs Pro again.' DeepInfra already serves it at fp8.
Developers praised the combination of license and context length as a cheap option for long agent loops. It competes directly with OpenAI's newly cut cached-input pricing on GPT-6.1 Sol. DeepSeek also put DeepSeek Harness, a desktop app for local workflows, into public preview on macOS, Windows and Linux.
One caveat is that DeepSWE results are self-reported across vendors, and real-world agent reliability often diverges from benchmark rankings. A Reddit thread on r/DeepSeek about 'never ending loop' behavior suggests rough edges remain. Watch for independent evaluations and for whether V4.1-Flash runs efficiently on Huawei Ascend hardware via TileLang.