TBPN

← Full issue

August 31, 2026

OpenAI’s reported Jalapeño inference chip targets higher efficiency

A report cited by Dylan Patel describes OpenAI’s Jalapeño as a data-center, rack-scale accelerator designed for large-language-model inference rather than training. It reportedly delivers roughly 1.5x to 1.9x more useful inference throughput per watt than NVIDIA’s GB200 and GB300 systems, while reducing end-to-end latency by roughly 1.7x to 3.6x.

OpenAI is reportedly planning a 10-gigawatt deployment of OpenAI-designed accelerator systems manufactured by or developed in partnership with Broadcom. Deployment is expected to begin in the second half of 2026 and continue through 2029; the agreement may cover second- or third-generation systems rather than exactly the Jalapeño chip. The planned capacity is described as roughly five times the compute currently operated by the company.

Privacy ·