Back
NVIDIASeptember 11, 20262 sources

NVIDIA releases free 30B Nemotron 3.5 Lightning, NIM 2.0.12 boosts throughput 2.5x

AI Analysis

Nemotron 3.5 Lightning is aimed squarely at the local-inference community: a free 30B model that runs 4x faster than predecessors and fits on a single RTX GPU, deployable through the popular Ollama, LM Studio, and llama.cpp stacks, or on cloud data-center shapes. That combination — free weights, single-GPU footprint, mainstream tooling — targets developers avoiding cloud bills, a segment that rallied around it in r/LocalLLaMA.

On the enterprise side, NIM 2.0.12 delivers 2.5x throughput on Nemotron 3 Ultra, meaning the same GPU fleet can serve more concurrent users — a software-only efficiency gain that stretches scarce hardware further, echoing the week's capacity theme from a different angle.

Strategically, NVIDIA giving away capable open-weight models serves its hardware business: every developer running Nemotron locally is a reason to buy NVIDIA GPUs, and every enterprise squeezing more users per GPU via NIM is deeper into the CUDA ecosystem. It also positions NVIDIA against Meta's Muse Spark pivot away from open weights and Alibaba's Qwen releases — filling the open-model gap Meta's Llama retreat left, while ensuring the models run best on NVIDIA silicon.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog