Back to Newsroom

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

By Modelverse Editorial·August 25, 2026·2 min read
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

OpenAI unveiled details of its Jalapeño inference accelerator at the Hot Chips conference, presenting the first benchmark results from Semianalysis’s InferenceX suite. The chip delivered higher tokens per user and greater throughput per kilowatt than the current state‑of‑the‑art inference processors, which the company compared against an Nvidia Blackwell system. OpenAI’s head of hardware, Richard Ho, noted that Jalapeño can serve more AI work per unit of power while returning responses with lower latency, and he projected initial deployment in small volumes by the end of 2026, scaling to broader use in 2027.

Jalapeño was co‑developed with Broadcom, leveraging OpenAI’s own models to guide the design process. The chip is positioned as a multigenerational platform, intended to evolve alongside future AI products, models, and memory subsystems. By adopting a full‑stack approach, OpenAI targeted specific stages of the inference pipeline that commonly create bottlenecks—particularly the prefill phase and the inter‑chip communication steps. The architecture minimizes data movement and keeps the KV cache local, allowing the system to activate the optimal combination of compute, memory, and networking for each inference stage.

  • Collaboration with Broadcom for silicon design and integration
  • Use of OpenAI’s internal models to inform architectural choices
  • Focus on reducing prefill and communication latency via localized KV cache
  • Planned multigenerational rollout beginning late 2026 with expanded deployment in 2027

Why this matters

Based on the reported improvements in tokens per user and power‑efficiency, one can infer that Jalapeño could lower operating costs for large‑scale inference workloads if the performance gains translate to real‑world deployments. However, the timeline places the chip’s arrival after several expected advances in competing accelerators, so its actual advantage will depend on how well OpenAI’s full‑stack optimizations keep pace with broader industry trends in process technology and memory bandwidth.

Share this article

Found this insightful? Share it with your community on Reddit, X, or copy the link.

ai-newsbrieftechcrunch-ai

Footnotes & Primary References

Related content

Accel-backed Keenable is indexing the web for AI agents

Now exiting stealth mode with a $26 million seed round, Keenable has been building a vast web search index for AI agents.

Read article

'The world seems to be ready': An interview with OpenAI head of product Thibault Sottiaux

TechCrunch talks agents, UX, and reporting to Greg Brockman with OpenAI's head of product.

Read article

Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark

NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technologies and increased v...

Read article