DeepSeek releases V4.1-Flash with 1M-token context, plans ~$75B Shanghai IPO

Per Reuters, SiliconAngle and The Register, V4.1-Flash is the smallest entry in a new architecture family but claims to beat the larger V4-Pro. The design's headline trick is asymmetry: it activates far less compute for reading (input) than for generating (output), and shrinks the KV cache to ~890 bytes per token — roughly 4x smaller than V4-Flash — enabling a native 1M-token window and 400+ tokens/sec throughput. Pricing starts around $0.003 per million cached input tokens, extending DeepSeek's cost-disruption playbook.
Parameter counts became the community flashpoint. DeepSeek cites 552B total (8B active), but r/LocalLLaMA users argued the real figure is 748B ('748B, not 552B,' 309 upvotes), and one source pegged it at 763B — reflecting persistent confusion over how MoE totals are counted. The top HN comment warned the low active-parameter figure is 'a marketing trick—you still need the VRAM to load weights.' Praise centered on long-context handling and tool-calling rather than raw MMLU, with friction over DeepSeek's custom license versus Apache 2.0.
The launch dominated r/LocalLLaMA (a thread hit 1,538 upvotes) and r/DeepSeek, where users reported one-shot coding wins for $0.07 and cheered DeepSeek auto-switching users to the cheaper-but-better model — contrasted favorably against 'greedy US closed' labs. It also lands amid the U.S. and Anthropic accusations that DeepSeek distilled Western models, a cloud over its rise.
Financially, DeepSeek retained CITIC Securities for a domestic STAR Market listing at roughly $75B, signaling ambitions to fund frontier compute at scale. Watch independent benchmarks, the real memory footprint, and whether the distillation allegations complicate its IPO or international availability.