OpenAI and Broadcom unveil LLM inference chip
OpenAI Jalapeño boosts LLM inference with better performance per watt See how the new Broadcom chip could speed AI services and cut costs
OpenAI and Broadcom have introduced Jalapeño, a new AI accelerator designed specifically for large language model inference. The companies said the chip is built to improve performance per watt and support current and future LLM workloads across a multigeneration compute platform.
According to OpenAI, Jalapeño was developed from design to tapeout in nine months with help from OpenAI models, Broadcom’s silicon expertise, and partners including Celestica. The chip is intended to reduce data movement, improve efficiency, and support largescale deployment in data centers.
OpenAI said early testing suggests the accelerator performs substantially better on energy efficiency than current stateoftheart systems, though full technical results have not yet been released. The first deployments are expected to begin in 2026 as part of a broader infrastructure effort aimed at making AI services faster, more reliable, and more affordable.