OpenAI's self-designed Jalapeño chip reportedly beats Nvidia Blackwell on efficiency

OpenAI revealed benchmark results for Jalapeño, its first in-house inference chip, designed from scratch with Broadcom specifically for large-language-model inference. According to the results, Jalapeño outperformed Nvidia's GB200 and GB300 Blackwell superchips on an inference benchmark, delivering higher throughput, lower latency, and strong energy efficiency. Design work began in mid-2024 and reached tape-out in roughly 16 months, an unusually fast timeline for a first-generation ASIC. Sam Altman called it 'very fast' on X.
The strategic significance is OpenAI reducing its dependence on Nvidia for the most cost-sensitive workload — inference, which scales with usage. A credible in-house inference ASIC would improve unit economics for ad-supported free tiers and high-volume products, and give OpenAI leverage in compute negotiations.
Competitively it intensifies the theme of Nvidia's moat shifting from chips toward capital and ecosystem: even as Nvidia backs tens of billions in infrastructure financing and reportedly moves to acquire Hugging Face, its silicon advantage on inference is being challenged. Google (TPUs), Amazon (Trainium/Inferentia), and Meta already build custom accelerators.
Skeptics urge caution. SemiAnalysis noted first-generation ASICs are rarely competitive out of the gate, making a claimed Blackwell-beating result in ~16 months remarkable and worth independent scrutiny. Benchmarks are also workload-specific — beating Blackwell on one inference test doesn't equal broad superiority, and software/ecosystem maturity favors Nvidia. Watch for third-party benchmarks, volume/yield details, and whether Jalapeño actually deploys across OpenAI's fleet.