OpenAI unveiled details of its Jalapeño inference accelerator at the Hot Chips conference, presenting the first benchmark results from Semianalysis’s InferenceX suite. The chip delivered higher tokens per user and greater throughput per kilowatt than the current state‑of‑the‑art inference processors, which the company compared against an Nvidia Blackwell system. OpenAI’s head of hardware, Richard Ho, noted that Jalapeño can serve more AI work per unit of power while returning responses with lower latency, and he projected initial deployment in small volumes by the end of 2026, scaling to broader use in 2027.
Jalapeño was co‑developed with Broadcom, leveraging OpenAI’s own models to guide the design process. The chip is positioned as a multigenerational platform, intended to evolve alongside future AI products, models, and memory subsystems. By adopting a full‑stack approach, OpenAI targeted specific stages of the inference pipeline that commonly create bottlenecks—particularly the prefill phase and the inter‑chip communication steps. The architecture minimizes data movement and keeps the KV cache local, allowing the system to activate the optimal combination of compute, memory, and networking for each inference stage.
- Collaboration with Broadcom for silicon design and integration
- Use of OpenAI’s internal models to inform architectural choices
- Focus on reducing prefill and communication latency via localized KV cache
- Planned multigenerational rollout beginning late 2026 with expanded deployment in 2027
Why this matters
Based on the reported improvements in tokens per user and power‑efficiency, one can infer that Jalapeño could lower operating costs for large‑scale inference workloads if the performance gains translate to real‑world deployments. However, the timeline places the chip’s arrival after several expected advances in competing accelerators, so its actual advantage will depend on how well OpenAI’s full‑stack optimizations keep pace with broader industry trends in process technology and memory bandwidth.
Share this article
Found this insightful? Share it with your community on Reddit, X, or copy the link.
