Richard Ho, OpenAI’s head of hardware, framed the new chip as a decisive leap in efficiency during a presentation at the Hot Chips conference. Unlike general-purpose hardware, Jalapeño is built as a multigenerational platform that synchronizes chip architecture with the specific demands of OpenAI’s own models. The system specifically targets the prefill and communication phases of inference, which are notorious for creating lag in high-demand environments.
In section Startups & Technology
OpenAI Unveils Jalapeño Chip Benchmarks for AI Inference
Jalapeño, the custom AI processor developed by OpenAI and Broadcom, outperformed Nvidia’s Blackwell system in initial InferenceX benchmarks. By prioritizing energy efficiency and high throughput, the hardware aims to solve critical data bottlenecks that typically slow down large-scale generative AI model responses.

By keeping the KV cache and model state local, the chip reduces the physical distance data must travel across the system. While current test results show superior tokens per watt compared to existing industry standards, the hardware faces a long road to market. OpenAI anticipates small-volume deployment by late 2026, with a broader rollout scheduled for 2027. By then, the company will have to contend with further iterations of Nvidia’s hardware, setting the stage for a high-stakes battle over data center efficiency.
Comments (0)
No comments yet. Be the first!