Close-up of a computer chip on a green circuit board with visible components and labeled text.

OpenAI Says Jalapeño Chip Outperforms Nvidia Blackwell in Inference Benchmarks

OpenAI says its Jalapeño chip outperforms Nvidia Blackwell in inference benchmarks and could begin limited deployment in late 2026.

In short

OpenAI says its custom Jalapeño inference chip outperformed a Nvidia Blackwell system on benchmark tests measuring speed and power efficiency. The chip is slated for very limited deployment at the end of 2026, with wider rollout expected in 2027.

  • OpenAI says Jalapeño beat current top inference hardware on speed and efficiency benchmarks.
  • The chip is designed to reduce bottlenecks in prefill, communication and KV-cache handling.
  • OpenAI plans a limited deployment in late 2026, with larger rollout in 2027.
  • The project is a collaboration with Broadcom and part of OpenAI’s long-term hardware strategy.

OpenAI has published its first benchmark results for Jalapeño, its custom inference chip, and says the new system delivers more tokens per user and higher throughput per watt than leading processors currently on the market. The company shared the findings at the Hot Chips conference on Tuesday, signaling that the chip is moving closer to a limited rollout at the end of 2026 and broader deployment in 2027.

The results matter because they suggest OpenAI is pushing beyond software into the hardware stack that powers its models, with the goal of making AI responses faster and cheaper to serve at scale. If the numbers hold up in real-world deployment, Jalapeño could help OpenAI reduce its dependence on third-party accelerators and improve the economics of inference, the compute-heavy process that happens every time a model generates an answer.

What OpenAI revealed at Hot Chips

OpenAI used the annual chip-design conference to give a closer look at Jalapeño, a system the company says was built specifically for inference rather than training. The update included benchmark data that OpenAI says shows a meaningful jump over state-of-the-art inference hardware available today.

According to the company, Jalapeño was tested on Semianalysis’s InferenceX benchmark and came out ahead on two measures that matter to AI operators: throughput per kilowatt and tokens per user. In practical terms, that means the chip can process more work for the same amount of power while also producing responses faster for individual users.

Richard Ho, OpenAI’s head of hardware, said in a press call that the benchmark results represent a major step forward relative to existing systems. He framed the chip as efficient enough to serve large numbers of users while still keeping latency low.

OpenAI’s hardware chief said the benchmark results show a substantial improvement over the current state of the art, with the chip designed to serve more AI work per unit of power and return answers more quickly.

Why inference hardware has become the new battleground

Inference is where the economics of modern AI get real. Training a large model is expensive, but serving that model to millions of people can be even more operationally demanding because every query requires memory access, communication, and compute coordination in real time.

That makes inference efficiency a critical advantage for any company running consumer or enterprise AI products at scale. The better a chip performs on latency and power draw, the more queries it can handle before costs rise too sharply.

OpenAI’s move suggests it sees hardware design as a strategic lever, not just a procurement issue. By tailoring the system to its own models and workloads, the company is trying to optimize the full path from prompt to response rather than relying on general-purpose accelerators alone.

How does Jalapeño differ from a standard chip strategy?

It differs because OpenAI is designing the chip in concert with the rest of its stack. Instead of treating compute, memory, networking, and models as separate layers, the company says it is building them together so the system can be tuned for the specific phases of inference.

That approach can reduce wasted movement of data and remove bottlenecks that show up when model state has to travel between components. It also gives OpenAI more control over how its products are delivered, from internal model behavior to the physical infrastructure behind them.

How Jalapeño is designed to speed up responses

OpenAI says Jalapeño focuses on the phases of inference that often create delays, especially prefill and communication. Prefill is the stage where a model processes the user’s input before generating output, while communication includes the back-and-forth movement of data between compute elements.

The company says those stages often become choke points in large-scale deployments. Jalapeño is meant to reduce those delays by minimizing data movement and keeping key model state local when possible.

That includes the KV cache, a memory structure used while a model is generating a response. OpenAI says the chip lets that cache be placed explicitly and kept close to where it is needed, while the system activates the right mix of compute, memory, and networking for each step.

This type of design is especially important for systems that need to serve huge numbers of requests at once. Lower latency improves the user experience, but it can also improve capacity because faster turnaround means hardware can be reused more efficiently across workloads.

How strong are the benchmark claims?

OpenAI’s benchmark comparison is notable, but it comes with an important caveat: the reference point is a Nvidia Blackwell system, and hardware competition in AI is moving quickly. By the time Jalapeño is widely deployed, rival chips and system designs may have advanced further.

That timing matters because benchmark leadership at announcement time does not guarantee a durable edge. The AI hardware market is highly dynamic, with vendors racing to improve memory bandwidth, interconnects, power efficiency, and specialized inference performance.

Even so, OpenAI’s results show that it believes its chip can compete with the best inference hardware in the market. For a company that has largely depended on outside silicon suppliers, that is an important public statement of intent.

Item Details Why it matters
Chip name Jalapeño OpenAI’s custom inference system
Public reveal Hot Chips conference, Tuesday First detailed benchmark discussion
Benchmark used Semianalysis InferenceX Measures inference performance and efficiency
Reported strengths More tokens per user; higher throughput per kilowatt Signals better latency and power efficiency
Competitor comparison Nvidia Blackwell system Sets the hardware bar Jalapeño is targeting
Expected rollout Small volumes at end of 2026; larger deployment in 2027 Shows the chip is still in an early commercialization phase

What is OpenAI’s multigenerational hardware plan?

OpenAI says Jalapeño is not intended to be a one-off experiment. Instead, it is being developed as part of a multigenerational platform in which AI products, models, chips and memory are designed together over time.

That kind of integrated plan could give the company more room to optimize each layer of the stack for its own workloads. In effect, OpenAI would be shaping both the software intelligence and the physical machinery that runs it.

This is a significant strategic shift for a company best known for models like ChatGPT. It places OpenAI in a smaller, more specialized club of technology companies that are not just buying infrastructure but trying to define it.

Why Broadcom is part of the picture

OpenAI says Jalapeño was developed closely with Broadcom, and the collaboration is central to how the chip came together. Partnering with a major semiconductor company gives OpenAI access to manufacturing and design expertise while letting it concentrate on model-aware optimization.

The company also said its own models helped during development, a detail that suggests the chip was shaped using AI-assisted design and validation. That kind of feedback loop could become increasingly important as hardware is tuned for model behavior rather than generic workloads.

When will Jalapeño actually reach customers?

OpenAI says the first deployment is expected at the end of 2026, but only in very small volumes. Broader deployment is expected in 2027 if development and production stay on track.

That timeline suggests the chip is still well short of full commercial scale. For now, the benchmark results are best understood as an early indicator of OpenAI’s direction rather than proof of a finished product ready to reshape the market.

Still, even a limited rollout could be meaningful if OpenAI uses Jalapeño to offload part of its inference demand. The company’s AI services are widely used, and any reduction in cost per token could matter across massive traffic volumes.

What this means for Nvidia and the wider AI chip race

OpenAI’s benchmark win, if sustained, would add more pressure to the market leader in AI accelerators. Nvidia remains the dominant supplier in the sector, but major customers have been looking for ways to diversify supply, lower costs, and reduce dependence on a single vendor.

Custom silicon has become one of the clearest responses to those pressures. Cloud providers, platform companies, and AI labs are all trying to carve out specific advantages by designing chips around their own workloads.

For OpenAI, the incentive is especially strong. Its products depend on rapid and efficient inference at massive scale, and even modest improvements in latency or power efficiency can translate into meaningful operational savings.

At the same time, the challenge is formidable. The leading chip vendors have deep experience, vast manufacturing partnerships, and a long track record of iterating quickly. OpenAI’s promise will depend not only on benchmark results but on whether it can bring the system into production reliably and at useful scale.

What should investors and industry watchers look for next?

They should watch three things: manufacturing progress, deployment volume, and whether OpenAI publishes additional real-world performance data. Those factors will reveal whether Jalapeño is an internal optimization tool or a genuine platform shift.

Observers will also want to see whether the chip’s early benchmark edge survives broader deployment, when software stacks, traffic patterns, and thermal constraints become harder to control. Inference hardware often looks different outside the lab.

Timeline of Jalapeño’s rollout

Here is the sequence of events that defines OpenAI’s latest hardware push:

  1. October 2025: OpenAI first announced Jalapeño.
  2. Tuesday at Hot Chips 2026: The company presented more technical detail and benchmark results.
  3. End of 2026: OpenAI expects the chip to enter very limited deployment.
  4. 2027: Broader adoption is expected if development stays on schedule.

That staged approach is typical for advanced semiconductor projects, especially when the hardware is being built for a highly specific workload. The road from prototype to production usually includes validation, packaging, supply chain coordination and system integration.

Why this matters beyond OpenAI

OpenAI’s Jalapeño program highlights how the AI race is moving from model quality alone to the full stack that supports it. The companies best positioned for the next phase may be those that can align software, hardware, networking and memory around a single performance goal.

That shift could influence how future AI systems are built across the industry. If custom inference chips prove worthwhile, more model developers may seek similar control over the infrastructure beneath their products.

It also underscores a broader reality: the cost and speed of AI are now central product questions, not just engineering details. For providers serving millions of users, inference efficiency can shape everything from margins to response times to the practicality of new features.

OpenAI’s latest update does not settle those questions, but it does show the company is betting that hardware will be a decisive part of its future. Jalapeño is a signal that the AI competition is extending deeper into silicon, where every watt, cache access and memory transfer can become a competitive advantage.

For now, the headline is simple: OpenAI says Jalapeño is faster and more efficient than current state-of-the-art inference hardware in benchmark testing, and the company plans to begin limited deployment by the end of 2026.

Frequently asked questions

What is OpenAI’s Jalapeño chip?

OpenAI’s Jalapeño chip is a custom hardware system built for AI inference, the stage where models generate responses to user prompts. The company says it is optimized for speed, efficiency and lower-latency serving at scale.

Did Jalapeño really outperform Nvidia Blackwell?

OpenAI says Jalapeño outperformed a Nvidia Blackwell system on the Semianalysis InferenceX benchmark, with better throughput per kilowatt and more tokens per user. That result is benchmark-based, so real-world performance will need further validation.

When will Jalapeño be available?

OpenAI says Jalapeño should appear in very small volumes at the end of 2026, with more significant deployment expected in 2027. The timeline suggests the chip is still in an early rollout phase.

Why is OpenAI building its own chip?

OpenAI is building its own chip to improve inference efficiency, lower power use and reduce delays in serving AI responses. A custom chip also gives the company more control over the full technology stack behind its products.

Share this 🚀