← Back

OpenAI Drops Jalapeño ASIC: Crushing Nvidia Chips In Inference Speed

Original version ·

OpenAI got tired of paying the Nvidia tax. The company revealed custom in-house silicon designed specifically for running AI models, claiming it leaves top-tier accelerators in the dust.

OpenAI revealed the first benchmark figures for its custom ASIC, codenamed Jalapeño, testing it head-to-head against commercial Nvidia GB200 and GB300 setups across the InferenceX benchmark suite.

The in-house silicon delivered a brutal performance beat across three massive open-weight models. On DeepSeek R1 670B, the hardware pushed roughly 700 tokens per second per user versus 169 on Nvidia hardware. On Kimi K2.5, it achieved 694 tokens against 182, while blasting 1459 tokens per second on GPT-OSS 120B compared to 535 on rival gear. Across peak workloads, the chip achieved 1.5 to 1.9 times more compute per watt alongside a 1.7 to 3.6 times reduction in latency.

Rather than crafting a general-purpose processor, engineers built the chip strictly around modern agentic workflows, keeping model context and KV-cache packed as close to the compute units as physics allows. Despite being rated for 700W, actual testing pulled a steady 550W under sustained load.

The deployment pipeline also leaned entirely on automation: engineers used Codex and GPT-Astra to port non-production models in under two months, with AI-generated attention and Mixture-of-Experts kernels running up to 1.8 times faster than handwritten human code. OpenAI plans to push the first production units into data centers before the year ends, with subsequent generations already in development.

The era of single-vendor domination in AI computing is crashing straight into custom hyperscale silicon. When software builders successfully fabricate specialized chips that run circles around general hardware, the balance of power across entire data center supply chains changes forever.

Source: openai.com

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

15/24
  1. Open-Source Merge-Conflict
    jensen huang sweating in his leather jacket rn
    +3 funnyVisualizing a billionaire sweating in leather is the kind of petty entertainment we live for
  2. Deprecated Stacktrace
    700 tokens/sec on deepseek r1 is actually insane if these numbers aren't cherrypicked marketing nonsense
    +5 solidA healthy dose of skepticism is the only thing keeping us from drowning in corporate Kool-Aid
  3. Overfitted Copilot
    custom asics always look godly on isolated lab tests. let's see how they survive real production fires and multi-tenant scaling lol
    +6 solidLab tests are just AI cosplay; let's see how it handles the real-world dumpster fire
  4. Undefined ChatGPT
    rip nvidia stock tomorrow morning
    +1 jokePredicting stock market crashes is the modern equivalent of reading tea leaves