Back
DeepSeekSeptember 14, 20262 sources

DeepSeek Releases MIT-Licensed V4.1-Flash at $0.15/M Tokens with 1M Context

AI Analysis

V4.1-Flash continues DeepSeek's pattern of shipping frontier-competitive open weights at a fraction of Western pricing. The 552B-parameter mixture-of-experts model activates only 8B–16B parameters per token, includes native visual understanding, and offers a 1M-token context window — all under a permissive MIT license at $0.15/M input and $0.60/M output tokens. Benchmarks circulated by DeepSeek and third parties claim it matches or beats Claude Opus 5 and GPT-5.6 Sol on several agentic coding tasks.

The key technical claim is a 75% reduction in KV-cache HBM requirements, which materially lowers the memory cost of long-context inference and is a big deal for anyone self-hosting. That efficiency led DeepSeek to phase out its older V4 Pro model in favor of Flash on both performance and cost grounds.

The open-weight release had immediate consequences: within hours of the September 10 push to Hugging Face, developers stripped the safety layers and uploaded 'abliterated' and 'uncensored' forks, with at least two community builds live and drawing download traffic by September 13. On r/LocalLLaMA the enthusiasm came with a caveat — local runs showed baked-in political censorship persists, prompting the line 'open weights ≠ open model.'

Corporate context sharpens the picture: DeepSeek hired a dealmaker CFO, per Reuters, ahead of a possible Shanghai STAR Market IPO, signaling it intends to institutionalize. The competitive implication is stark — a Chinese lab is setting the price-performance frontier for open agentic models, pressuring both proprietary incumbents and the economics of closed APIs. Watch whether the censorship and provenance concerns limit Western enterprise adoption.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog