In short
OpenAI says its in-house chip, Jalapeño, outperformed Nvidia’s latest superchips on AI inference benchmarks. The company plans a small rollout by the end of 2026 and a broader deployment in 2027.
- OpenAI says Jalapeño beats Nvidia GB200 and GB300 systems on key inference metrics.
- The chip is built as an ASIC with Broadcom and is focused on AI inference, not training.
- OpenAI reported 1.5x to 1.9x better work per watt and much lower latency in tests.
- The company plans limited deployment by late 2026, with scaling expected in 2027.
- OpenAI says it will still rely on partners like Nvidia rather than replacing its entire hardware stack.
OpenAI says its new in-house chip, Jalapeño, can deliver faster responses and better efficiency than rival systems in AI inference, a step that could reduce the company’s dependence on outside suppliers as demand for its models keeps rising. The chip, built with Broadcom, is expected to ship in small numbers by the end of 2026 before OpenAI scales deployment in 2027.
The announcement matters because it puts OpenAI more directly into the AI hardware race at a moment when compute capacity has become one of the most important constraints in the industry. By pursuing its own inference chip, OpenAI is trying to improve performance, control costs, and secure more reliable access to the hardware that powers real-time AI products and agents.
What OpenAI says Jalapeño can do
OpenAI’s headline claim is that Jalapeño outperformed the strongest Nvidia systems it compared against on an AI inference benchmark. In a blog post and follow-up briefing, the company said the chip offered both lower latency and higher throughput, two metrics that AI infrastructure teams often have to balance against each other.
Lower latency means a system can return an answer more quickly. Higher throughput means it can handle more work overall. OpenAI hardware vice president Richard Ho described Jalapeño as offering the “best of both worlds,” arguing that many AI systems must sacrifice one advantage to improve the other.
According to OpenAI hardware vice president Richard Ho, most AI systems face a trade-off between speed and overall work capacity, while Jalapeño is meant to improve both at once.
OpenAI says the chip completed between 1.5 and 1.9 times more AI work per watt than the comparison systems across three benchmarked models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. The company also said Jalapeño delivered end-to-end latency that was 1.7 to 3.6 times lower than those Nvidia-based comparisons.
Why these benchmark results matter
These measurements matter because they speak directly to the economics of running large AI products at scale. A chip that does more useful work for each watt of power can lower operating costs, reduce heat, and make it easier to deploy more inference capacity in the same data-center footprint.
For users, the practical result could be quicker replies from chatbots, more responsive AI agents, and steadier service as traffic grows. That is especially relevant for OpenAI, whose products depend on serving millions of requests efficiently and consistently.
How Jalapeño fits into OpenAI’s broader hardware plan
Jalapeño is not a general-purpose server chip. OpenAI says it is an application-specific integrated circuit, or ASIC, designed for AI inference, the stage where a trained model is used to generate an answer, complete a task, or power an agent in production.
The chip was first revealed in June and was developed in partnership with Broadcom. That makes it part of a growing industry trend in which major AI companies design custom silicon to optimize the workloads that matter most to them rather than relying entirely on standard accelerators.
OpenAI’s strategy appears to be additive rather than replacement-based. Ho said the company does not intend to swap out its full chip lineup for Jalapeño. Instead, OpenAI wants to keep working with existing suppliers, including Nvidia, while adding its own silicon into the mix.
| Item | Details | Why it matters |
|---|---|---|
| Chip name | Jalapeño | OpenAI’s first publicly discussed in-house inference chip |
| Chip type | ASIC | Built for a specific workload rather than general computing |
| Partner | Broadcom | Provides key hardware-development support |
| Benchmark platform | InferenceX | Used to compare inference performance across systems |
| Reported advantage | 1.5x to 1.9x more work per watt | Signals improved efficiency |
| Latency improvement | 1.7x to 3.6x lower latency | Suggests faster end-user responses |
| Deployment timeline | Small volumes by end of 2026; ramp in 2027 | Shows the chip is still early in rollout |
Why OpenAI is building its own chip now
The answer is simple: AI demand is growing faster than the infrastructure that supports it. Training and serving large models requires enormous amounts of compute, and the companies that can secure efficient hardware have a major advantage in cost, speed, and reliability.
For OpenAI, custom chips could help in several ways:
- reduce dependence on a small group of hardware vendors
- improve the efficiency of model serving
- lower latency for interactive products
- support AI agents that need rapid, repeated inference
- create more predictable access to compute as usage grows
The timing is also notable because Nvidia’s flagship systems remain dominant in the market. A competitive inference chip from OpenAI would not immediately displace that ecosystem, but it could give the company leverage as it negotiates for future capacity and designs more tightly optimized systems for its own workloads.
How inference differs from training
Inference is the process of using a finished model, while training is the expensive phase where the model learns from data. Inference is what users experience when they ask a chatbot a question or ask an agent to complete a task, making it the part of AI that has the most direct consumer and enterprise impact.
Because inference happens continuously after a model launches, improvements in latency and efficiency can produce major long-term savings. That is why many companies are increasingly building or buying chips tailored specifically for serving models rather than only for training them.
How does Jalapeño compare with Nvidia’s latest systems?
OpenAI’s benchmark placed Jalapeño against Nvidia’s GB200 and GB300 superchips, which the company described as the strongest results available at the time. The comparison was designed to show how OpenAI’s custom hardware stacks up against the most advanced widely deployed alternatives in the market.
According to OpenAI, Jalapeño came out ahead on both efficiency and latency in the benchmark results across the three models tested. The company did not publish every underlying hardware detail in the source material, but the broad claim is that the chip can do more useful inference work while drawing less power and taking less time to respond.
That said, benchmarks are not the same as broad real-world deployment. Performance can vary by model, workload, memory configuration, software stack, and system integration. OpenAI’s results are important, but they should be viewed as an early indicator rather than the final word on the chip’s value.
What the benchmark numbers suggest
The reported numbers suggest that OpenAI is targeting workloads where milliseconds matter and power consumption is a major operating concern. If those figures hold up in broader production use, the company could meaningfully improve the economics of serving high-volume AI requests.
That advantage would be especially valuable for products that depend on immediate, back-and-forth interaction, including chatbots, coding assistants, and more autonomous AI agents that may need to make repeated calls to models in a short period of time.
What happens next for OpenAI’s chip rollout?
OpenAI says Jalapeño will reach the company in small quantities by the end of this year, with volume expected to increase in 2027. The company did not disclose how many chips it plans to deploy in 2026, leaving the scale of the rollout unclear.
That cautious rollout suggests OpenAI is still in the early phase of integrating custom silicon into its infrastructure. The company is likely to spend the next several quarters validating performance, tuning software, and testing how the chips behave under live demand before committing to a much larger deployment.
OpenAI also said it will keep developing second- and third-generation versions of the chip. That signals a multi-year hardware roadmap, not a one-off experiment. If the company continues down this path, it could eventually have a family of proprietary chips built specifically around its own model-serving needs.
| Milestone | Timing | Expected significance |
|---|---|---|
| Public introduction | June 2026 | First confirmation of OpenAI’s ASIC effort |
| Benchmark disclosure | August 2026 | Claims of stronger inference performance vs. Nvidia systems |
| Initial deployment | By end of 2026 | Small-volume rollout inside OpenAI |
| Scaling phase | 2027 | Broader integration if early results hold up |
| Next chip generations | Ongoing | Signals long-term commitment to custom hardware |
Why the AI chip race is heating up
OpenAI’s move reflects a broader industry shift. The biggest AI players are no longer only competing on model quality; they are competing on the full stack, including power, networking, memory, software optimization, and the silicon that makes everything run.
That is why custom hardware has become a strategic priority for so many companies. A faster model is valuable, but a faster model that can be served more cheaply and reliably at scale is often even more important. The company that controls more of that stack has greater room to improve product performance and manage costs.
For OpenAI, Jalapeño could also be a hedge against supply constraints. The AI boom has repeatedly strained access to cutting-edge accelerators, and companies that depend entirely on external chip makers can find themselves at the mercy of manufacturing timelines and allocation decisions.
What this means for Nvidia
This does not mean Nvidia is losing its role in AI infrastructure. OpenAI explicitly said it still views Nvidia as an important partner, and the company plans to continue using a mixed hardware strategy. But custom chips from major customers do create a longer-term competitive threat if they gradually reduce the volume of third-party hardware a company needs to buy.
Even if Jalapeño only handles a slice of OpenAI’s workloads, it could help the company diversify its infrastructure and negotiate from a stronger position. In the AI hardware market, that kind of optionality is increasingly valuable.
How should the benchmark claims be read?
The answer is carefully: as a promising early result, not a full verdict. Benchmark tests can highlight real gains, but they are also shaped by the exact models, configurations, and measurement methods used. OpenAI’s figures indicate that Jalapeño is competitive on paper and potentially powerful in practice, but broader adoption will determine its true impact.
That distinction matters because the AI hardware field is full of partial wins. A chip may excel in a narrow workload while underperforming in another. OpenAI’s challenge will be proving that Jalapeño can sustain its claimed advantages across a wide range of production conditions, not just in controlled comparisons.
Still, the company’s confidence is notable. OpenAI would not likely emphasize inference efficiency against Nvidia’s top systems unless it believed the results represented a meaningful technical milestone. The rollout plan also suggests the company sees enough promise to begin operational deployment rather than keep the chip locked in testing.
The bigger strategic picture for OpenAI
Jalapeño is part of a larger effort to turn OpenAI from a software company that relies on external infrastructure into one that can shape more of the stack beneath its models. That shift could influence everything from service quality to product margins to the speed at which the company can launch more advanced agentic systems.
If future generations of the chip keep improving, OpenAI could eventually tune hardware and software together more tightly than is possible when buying off-the-shelf accelerators alone. That level of vertical integration has long been a competitive advantage in computing, and AI is now moving in that direction.
For now, though, the story is less about a sudden takeover and more about a strategic foothold. OpenAI’s Jalapeño chip appears to be an early but important attempt to make inference faster, cheaper, and more controllable — three goals that are becoming central to the AI industry’s next phase.
Key numbers at a glance
- 1.5x to 1.9x more AI work per watt, according to OpenAI
- 1.7x to 3.6x lower end-to-end latency versus comparison systems
- End of 2026 targeted for initial small-volume deployment
- 2027 expected to bring a broader ramp-up
- Three models used in the benchmark: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T
As OpenAI moves from model builder to hardware designer, Jalapeño offers a look at where the company thinks the AI market is heading: toward systems that are not only smarter, but faster, more efficient, and more tightly integrated from the silicon up.
Frequently asked questions
What is OpenAI’s Jalapeño chip?
OpenAI’s Jalapeño chip is a custom ASIC designed for AI inference, the stage where a trained model generates responses or completes tasks. OpenAI built it with Broadcom to improve speed, efficiency, and reliability for production AI systems.
Did Jalapeño really beat Nvidia’s latest chips?
OpenAI says yes, at least in its benchmark testing. The company reported that Jalapeño delivered more work per watt and lower latency than Nvidia GB200 and GB300 superchips across several model tests, though real-world performance could differ.
When will OpenAI use Jalapeño in production?
OpenAI says it expects to deploy Jalapeño in small volumes by the end of 2026. The company then plans to ramp up usage in 2027, although it has not said how many chips will be deployed next year.
Why does OpenAI want its own AI chip?
OpenAI wants more control over cost, performance, and supply. Custom hardware can reduce latency, improve power efficiency, and make it easier to scale AI products as demand rises, especially for chatbots and agents that need rapid inference.
Will OpenAI stop buying Nvidia chips?
No. OpenAI says it does not plan to replace its entire chip lineup with Jalapeño. The company says it will continue working with outside partners, including Nvidia, while adding custom chips into its broader compute strategy.









