Back
NVIDIASeptember 17, 20261 sources

NVIDIA releases 4-bit DeepSeek-V4.1-Flash-NVFP4 for Blackwell GPUs under MIT license

AI Analysis

NVIDIA released a 4-bit quantized version of DeepSeek's V4.1-Flash, branded NVFP4 and tuned specifically for its Blackwell GB300 GPUs. The optimized model supports the full 1-million-token context window without any retraining, and NVIDIA reports accuracy shifts of no more than 1.5 points across six benchmarks compared to the original — a small enough degradation to make the 4-bit version viable for production. It's distributed under the MIT license for both commercial and non-commercial use.

The technical significance is the NVFP4 format itself: 4-bit precision dramatically reduces memory footprint and increases inference throughput on Blackwell hardware, and doing so with under 1.5 points of benchmark drift on a 552B-parameter MoE model is a meaningful validation of NVIDIA's low-precision tooling. Combined with DeepSeek's already-tiny 890-byte KV cache, an NVFP4 build makes running a frontier-scale open model on a single high-end node far more practical.

Strategically, it's notable that NVIDIA is actively optimizing and publishing an official build of a Chinese open-weight model — a signal that NVIDIA's commercial interest is in maximizing GPU demand across whatever models developers want to run, regardless of origin. It also deepens the DeepSeek ecosystem momentum that has developers reporting subscription cancellations in favor of self-hosting.

The caveats: the accuracy figures are NVIDIA-reported on six benchmarks, and 4-bit quantization can degrade unpredictably on out-of-distribution or hard-reasoning tasks not captured by standard benchmarks — the same 'quality trade-off on harder reasoning' caveat DeepSeek users flagged for the base model. Still, for teams with Blackwell hardware, NVFP4 meaningfully lowers the cost of running a 1M-context frontier open model.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog