OpenAI just dropped Jalapeño: The chip designed to make Nvidia sweat
OpenAI and Broadcom finally unveiled Jalapeño, a custom ASIC for AI inference. It claims massive speed gains while using half the power of typical GPU setups. It is a bold move to cut out the middleman and challenge the Nvidia tax.
At the Hot Chips 2026 conference, OpenAI and Broadcom showcased Jalapeño, a custom inference ASIC that actually works. The hardware is a beast: a 128-chip rack pushing 1.7 exaflops at 4-bit precision with 2.7 PB/s of bandwidth. Unlike current industry trends that split prefill and decoding tasks across different systems, Jalapeño mashes them together to avoid idle time. It relies on a unique memory hierarchy where every core handles its own slice of HBM4, while sync is offloaded to a separate fabric.
The real secret sauce is software, not silicon. OpenAI built Gluon, a custom kernel language sitting on top of Triton, which allows them to map tensors across the chip layout with surgical precision. Using their own Codex model to write kernels, the team scaled from a test bench to a full rack in just eight days. This effectively bypasses the usual CUDA-lock argument, as the same team designing the hardware is writing the compiler, making the optimization loop significantly tighter.
However, this is not an instant Nvidia killer. The benchmarks focused on short context windows, conveniently ignoring the heavy-duty agentic workloads that actually drain compute budgets. While Broadcom is promising 50% lower costs per token compared to current GPUs, that is mostly about reclaiming the margin Nvidia currently takes, rather than some miraculous physics-defying breakthrough. First-generation chips are usually glorified paperweights, yet Jalapeño manages to be surprisingly competent at its narrow job.
The industry is watching whether OpenAI can actually scale this to thousands of nodes without the entire stack falling apart. If the biggest buyer on the planet stops buying from Jensen Huang, the entire supply chain logic shifts from 'can we get enough chips' to 'how low can we squeeze the price'. It is a classic move of a company deciding it is big enough to stop renting its future from someone else.
Source: The Register
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.