Back
NVIDIASeptember 11, 20262 sources

NVIDIA boosts Nemotron throughput 2.5x and expands media AI at IBC

AI Analysis

Per NVIDIA's blog and coverage, the Nemotron updates are about squeezing more inference out of installed hardware: NIM 2.0.12 claims a 2.5x throughput gain on Nemotron 3 Ultra, directly increasing how many concurrent users an enterprise can serve per GPU — an economically meaningful lever as demand outstrips supply industry-wide. Nemotron 3.5 Lightning, a 30-billion-parameter free model, is pitched as running 4x faster across local and cloud deployment shapes, extending NVIDIA's open-model footprint for developers.

The timing dovetails with NVIDIA's leaderboard push — Nemotron 3 Embed 8B ranked #1 for combined nDCG@10 on the Q2D-Web benchmark across 190M web documents and ~70K agent-reformulated queries in 10 languages, work NVIDIA credited to a Perplexity collaboration. Together these reinforce NVIDIA's strategy of not just selling chips but shipping competitive open models and serving software that make its hardware more productive.

At IBC (the broadcast industry show), NVIDIA demonstrated real-time broadcast AI, including a Synthetic Video Detector it says is now 99.3% accurate on text-based detection — a notable move into the deepfake/provenance space as synthetic media proliferates. That positions NVIDIA on both sides of the generative wave: enabling creation and offering detection.

Strategically, all of this lands the same week NVIDIA agreed to buy Hugging Face (a separate story) and touted a 2GW Australian AI-factory buildout and a Palantir sovereign-supply-chain collaboration — a full-court press to own the AI stack from silicon to models to distribution. Caveats: throughput and accuracy figures are vendor-reported and workload-dependent, and a 99.3% detector still fails on non-text-based synthesis. Watch independent Nemotron 3.5 Lightning benchmarks and real-world detector robustness.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog