OpenAI’s compute strategy treats the AI stack as a single, tightly coupled system where advances in silicon, software, models, and infrastructure reinforce each other. The company’s first custom inference accelerator, Jalapeño, was built to sit alongside existing third‑party accelerators and give OpenAI direct control over how its models are executed, measured in terms of throughput, latency, energy use, and cost. By co‑designing the chip with the serving stack, memory subsystem, and network fabric, OpenAI aims to shift the Pareto frontier for each workload class—frontier training, high‑volume inference, and always‑on agents—toward better performance per dollar.
On the public InferenceX benchmark using the GPT‑OSS 120B model, Jalapeño achieved higher peak throughput per kilowatt and lower token latency than the commercial systems it was compared against. The accelerator also showed strong results on DeepSeek R1 and Kimi K2, indicating that its efficiency gains are not limited to a single model family. OpenAI’s hardware portfolio now includes Microsoft Azure compute, NVIDIA GPUs, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank, allowing the company to match workloads to the strongest available solution while preserving economic flexibility.
Why this matters
The measured improvements in throughput per watt and latency demonstrate that a first‑party inference chip can deliver tangible efficiency advantages over existing market offerings when integrated with a purpose‑built software stack. This suggests that vertical integration—spanning chip design, model serving, and infrastructure—may become a viable path for AI providers seeking to optimize cost and performance at scale, rather than relying solely on heterogeneous third‑party accelerators. The ability to port those gains across diverse model families further underscores the architectural generality of the approach.
Share this article
Found this insightful? Share it with your community on Reddit, X, or copy the link.
