Tuesday, 25 August 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast

The Register ·
OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast

OpenAI offered its closest look yet at its spicy new Jalapeño AI accelerator at the annual Hot Chips semiconductor development conference at Stanford on Tuesday.

The chips, first teased earlier this year, were developed in collaboration with Broadcom, and are the first in a series of custom silicon from OpenAI, designed (in part) by AI, for AI.

Compared to contemporary GPU systems from Nvidia, OpenAI says the parts will deliver both higher throughput and lower latency when they start trickling out later this year and reach volume production in 2027.

To be clear, Jalapeño won’t replace OpenAI’s long-time hardware partners, which also happen to be some of its most important investors.

OpenAI still needs compute for training, and the highly programmable nature of GPUs means that OpenAI is likely to deploy on AMD and Nvidia first and then transition to in-house silicon later.

It’s also worth noting that AMD’s MI455X and Nvidia’s Rubin GPUs, also expected to ramp production in early 2027, are very different kinds of chips optimized for a mix of training and inference, whereas OpenAI’s custom silicon only needs to excel at one job: inference.

Memory bandwidth is king When it comes to inference, compute is key but memory bandwidth is king.

And based on early benchmarks OpenAI shared with the press before its Hot Chips presentation Tuesday, the chip is shaping up to be an inference beast.

Testing on SemiAnalysis’ InferenceX benchmark suite — presumably this is an unofficial test — shows OpenAI’s Jalapeño-based systems delivering between 1.5x and 1.9x more “AI work” at peak throughput, and 1.7x to 3.6x lower end-to-end latency than the competition across GPT-OSS-120B, DeepSeek R1, and Kimi K2.5.

If the latter two seem like weird models for OpenAI to be testing against, it's not that OpenAI plans to use these chips to run competitors' models, it’s just the models InferenceX uses.

In any case, the test shows that Jalapeño isn’t some model-specific architecture designed for maximum performance at the expense of programmability… cough, cough Taalas.

Meanwhile, for ultra-low-latency inference, which has become the hot new segment for AI infrastructure providers, OpenAI says its chips are 2.1x to 4.1x faster.

At a system level — we're starting here because the frontier models OpenAI trains rarely run on a single chip any more — each Jalapeño system with its 128 accelerators packs 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 and just shy of 2 petabytes a second of memory bandwidth.

By comparison, AMD and Nvidia’s latest rack systems are faster, delivering 1.46x to 2x more compute and up to 12 percent more memory on Helios, but just 85 percent the memory bandwidth of OpenAI’s rack.

Read the full article on The Register ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.

More from The Register

See all ›

More in Technology

See all ›