At Tuesday’s Hot Chips conference, OpenAI shared a more detailed picture of Jalapeno, including the first batch of benchmark results for the new system. When tested on SemiAnalysis’ InferenceX benchmark, Jalapeño recorded more tokens per user and throughput per kilowatt than any of the most advanced inference processors currently available.
“The bottom line is that the results demonstrate a very significant performance improvement over state-of-the-art technology,” Richard Ho, Head of Hardware at OpenAI, said in a press call. “Jalapeño can handle more AI work per unit of power while also returning responses more quickly. It’s very efficient to serve a large number of customers, but also with very low latency.”
It’s worth noting that this comparison is to an Nvidia Blackwell system. But by the time jalapeños are fully introduced, the competition may have advanced significantly. Ho predicted that Jalapenos would be deployed in “very small quantities” in late 2026, with a larger deployment in 2027.
First announced last October, Jalapeño was developed by OpenAI in close collaboration with Broadcom, and OpenAI’s proprietary models aided the development process. The company plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips, and memory to all be developed together.
Thanks to its full-stack approach, OpenAI was able to address specific phases of the inference process that often cause friction during the inference process. In particular, Jalapeño is designed to minimize delays in the prefill and communication phases of processing, which are often bottlenecks.
“We designed Jalapeno to minimize data movement and communication delays,” the company said in a blog post introducing the results. “This means that model state, including the KV cache used during response generation, can be explicitly placed and maintained locally while the system activates the appropriate combination of compute, memory, and networking during each inference phase.”
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
