Back
OpenAIJune 26, 20262 sources

OpenAI unveils Jalapeño, its first custom inference chip built with Broadcom

AI Analysis

OpenAI introduced Jalapeño, its first self-developed AI inference accelerator, co-designed with Broadcom and aimed at deployment by the end of 2026. Greg Brockman described it as 'designed from scratch for LLM inference over nine months, accelerated by our models,' adding that performance-per-watt is 'looking incredible' — a claim that, if it holds, would directly target the economics of running models at OpenAI's scale.

Inference accelerators are purpose-built for serving (not training) models, optimizing throughput, latency, and energy efficiency for the repetitive matrix work of generating tokens. By owning this layer, OpenAI joins Google (TPUs), Amazon (Trainium/Inferentia), and Meta in building in-house silicon to escape NVIDIA's pricing power and tailor hardware to its own kernels, memory, networking, and scheduling.

The competitive read was loud in the developer community: one widely shared take praised OpenAI for 'moving deeper into chips, kernels, memory, networking, racks, scheduling, deployment' as a serious full-stack hardware play. NVIDIA, for its part, is facing the same custom-chip pressure across big customers, even as Jensen Huang frames AI as a five-layer 'cake' that keeps overall compute demand soaring.

The caveats: Jalapeño is an announcement, not a shipping product, and a nine-month design timeline is aggressive for first silicon — yield, software maturity, and real-world perf-per-watt remain to be proven. Watch for benchmark disclosures, the Broadcom manufacturing details, and whether OpenAI commits volume that would actually dent its NVIDIA spend rather than merely diversify it.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog