In short
Nvidia is promoting its upcoming Vera Rubin system as a full-stack AI data center platform, not just a GPU upgrade. The company says the rack-scale design improves efficiency, simplifies deployment and strengthens its push into CPUs as competition with AMD intensifies.
- Nvidia is positioning Vera Rubin as a full AI data center system, with CPUs, GPUs and rack infrastructure.
- The company says the new platform delivers major gains in efficiency, memory bandwidth and deployment simplicity.
- OpenAI already has a Vera Rubin rack in use, while Microsoft and Oracle are early customers.
- Nvidia’s push comes just before AMD’s annual event, heightening the competitive stakes in AI hardware.
Nvidia is using a major pre-AMD product moment to argue that the next phase of artificial intelligence will require more than faster GPUs. The company says its upcoming Vera Rubin system will combine CPUs, GPUs, networking and cooling into a tightly integrated AI data center platform designed to power agentic workloads, improve efficiency and reduce installation time.
The pitch matters because Nvidia is trying to move from being the dominant supplier of AI accelerators to becoming the vendor that helps run nearly every part of an AI data center. That strategy could deepen its grip on the industry as rivals AMD and Intel push harder into the same market.
During a technical briefing at Nvidia’s Santa Clara headquarters last week, executives unveiled more details about Vera Rubin, the successor to Grace Blackwell, and framed it as a foundation for the company’s next growth phase. The system is expected to ship in the second half of this year, with early customers including Microsoft, OpenAI and Oracle, according to Nvidia.
What Nvidia is trying to sell now
Nvidia’s message is no longer limited to graphics processors. The company is increasingly presenting itself as a supplier of complete AI infrastructure, including the central processing units that manage data movement, networking and orchestration across modern AI systems.
That shift reflects a broader change in how large AI systems are built. Training still relies heavily on GPUs, but the rise of agentic AI — systems that break tasks into steps, interact with tools and move large amounts of data between components — has increased the importance of CPUs and system-level design.
In Nvidia’s view, the next generation of AI hardware must do more than accelerate math. It must also coordinate complex workflows efficiently enough to keep expensive chips busy and power use under control.
Ian Buck, Nvidia’s vice president of accelerated computing, told reporters the company is committed to shipping new architectures that include both GPUs and CPUs, arguing that constant innovation is essential to survival in Silicon Valley.
That broader ambition is central to Nvidia’s Vera Rubin push. Instead of selling only stand-alone chips, the company is emphasizing racks, systems and data center architecture as the real product.
How Vera Rubin differs from Grace Blackwell
Vera Rubin is Nvidia’s next major platform after Grace Blackwell, and the company is describing it as a more efficient, easier-to-deploy successor with far greater performance per watt. It is built around a one-to-two CPU-to-GPU ratio, with 36 Vera CPUs paired with 72 Rubin GPUs in a single NVL72 system.
The new platform is also meant to be simpler for customers to install and operate. Nvidia says the racks use far fewer cables than earlier generations and can be treated as “plug-and-play” compared with previous systems that required more complicated integration.
Nvidia executives described the new architecture as “cable-free compute” and “hot-swappable,” suggesting that customers could potentially move from a multi-hour rack installation process to one that takes only minutes.
| System | Key components | Main pitch | Deployment claim |
|---|---|---|---|
| Grace Blackwell | GPU-focused hybrid superchip | High-performance AI compute | Earlier generation, more complex integration |
| Vera Rubin NVL72 | 36 Vera CPUs + 72 Rubin GPUs | Full-stack AI infrastructure | Faster setup, fewer cables, hot-swappable racks |
| Standalone Vera CPU | ARM-based CPU | Orchestration and agentic workloads | Potentially available to some customers as soon as August |
The company says Vera Rubin is also designed to deliver much better performance efficiency than its predecessor, a key selling point as AI data centers become more power-hungry and expensive to operate.
Why Nvidia is emphasizing CPUs now
Nvidia is emphasizing CPUs because the AI industry is changing in ways that make central processors more valuable. As AI systems become more agentic, they need more coordination between memory, networking, storage and software tasks. That orchestration often falls to CPUs rather than GPUs.
This is a meaningful opportunity for Nvidia. The company has historically dominated AI training with GPUs, but there is still a large market for data center CPUs, especially as companies build larger and more complex AI clusters.
Nvidia’s CPU strategy is also a way to reduce its dependence on a single product category. By selling the CPU, GPU and rack as a package, it can capture more of the budget that hyperscalers and AI labs spend on infrastructure.
What is Nvidia saying about performance?
Nvidia says Vera Rubin NVL72 will process ten times as many tokens per watt as Grace Blackwell. The company also claims its Vera CPU handles agentic AI tasks faster than comparable CPUs from AMD and Intel.
Those comparisons, however, may not tell the whole story. Nvidia’s own tests reportedly used slightly older generations of rival processors, which means the benchmark claims should be viewed as directional rather than definitive. Even so, the company is clearly trying to position Vera Rubin as a meaningful leap in both throughput and efficiency.
Another major selling point is memory bandwidth. Nvidia says the new system’s localized memory subsystems offer nearly three times the memory bandwidth of Blackwell, a feature that could appeal to customers trying to work around shortages of high-bandwidth memory, one of the industry’s most constrained components.
For buyers building large clusters, that extra bandwidth could be just as important as raw compute power. If data can move faster between parts of the system, the AI application can spend less time waiting and more time processing.
How does the new hardware fit the AI data center race?
It fits into a larger contest to control the architecture of AI infrastructure, not just the chips inside it. Nvidia and AMD are now competing for the biggest multi-year contracts with hyperscalers such as Meta and Amazon, as well as AI labs including OpenAI, Anthropic and others.
These deals matter because AI buildouts are increasingly measured in rack-scale deployments, not single chips. Whoever sells the most complete system can potentially lock in software, networking and upgrade relationships for years.
Nvidia’s approach leans heavily on its own design choices. The company uses ARM-based CPUs for Vera Rubin, unlike the x86 architecture still dominant across most data center CPUs. AMD, by contrast, has built its reputation on x86 and chiplet-based designs that have become common in modern processors.
That difference gives the rivalry an architectural dimension. Nvidia is betting on a tightly integrated monolithic chip approach, while AMD has long championed chiplets as an efficient way to scale performance and manage manufacturing complexity.
Why does Nvidia dislike chiplets here?
Nvidia says chiplets impose a performance penalty in memory and data movement, which it describes as a kind of hidden cost on modern processors. By using a single monolithic chip for Vera Rubin, the company argues that data can move more quickly across one integrated circuit without the overhead of stitching together multiple pieces.
Hannah Coutand, who leads product marketing for Nvidia DGX Cloud, said during the briefing that interconnected chiplets create a “tax” on bandwidth and movement. Nvidia’s answer is to favor a more unified design that it says better suits AI workloads.
This is more than an engineering preference. It is a competitive statement about where Nvidia thinks the market is headed: toward platforms optimized for throughput and coordination rather than just isolated compute speed.
What makes Vera Rubin easier to deploy?
Nvidia says Vera Rubin is easier to deploy because it reduces physical complexity inside the rack. Less cabling means fewer connection points, less assembly work and fewer chances for installation errors.
The company also says the system is fully liquid-cooled, which matters because cooling is one of the biggest operating costs in modern AI data centers. Liquid cooling can be more efficient than air cooling and is especially useful when chips are packed densely into high-performance racks.
That combination of reduced cabling and improved cooling is intended to make the system more attractive to large buyers that want faster deployments and lower operational friction.
- Fewer cables can simplify rack assembly.
- Liquid cooling can reduce energy required for thermal management.
- Hot-swappable components can shorten maintenance windows.
- Integrated racks may lower total deployment complexity for hyperscalers.
Who is already using the new system?
Nvidia says OpenAI already has one Vera Rubin rack in use, an early sign that the company is moving its newest platform from lab concept to real-world deployment.
That kind of early adoption is strategically important. It gives Nvidia a marquee reference customer at a time when rivals are trying to prove their own next-generation AI systems can handle production workloads.
Microsoft and Oracle are also listed among the early customers Nvidia expects to support when the system ships later this year, though the company has not detailed the exact scope of those deployments.
Executives at the Santa Clara briefing portrayed Vera Rubin as a platform that is already moving into customer hands, not a speculative future product, while stressing that production ramp remains on track.
Why the timing matters ahead of AMD’s event
Nvidia’s disclosure wave comes just before AMD’s annual product event in San Francisco, where AMD is expected to showcase its own data center and AI chips. The timing is not accidental.
By releasing benchmarks and system details now, Nvidia is trying to shape the narrative before AMD’s event can capture attention. In the AI hardware market, perception matters almost as much as product specifications, especially when vendors are competing for long-term enterprise contracts.
AMD has been steadily gaining ground in data center CPUs, and its influence in the market has grown over the last two years. Nvidia is clearly aware that the competition is no longer hypothetical.
| Company | Core strength | Architecture emphasis | Current competitive angle |
|---|---|---|---|
| Nvidia | AI GPUs and rack-scale systems | ARM CPUs, monolithic design | Full-stack AI data center platform |
| AMD | Data center CPUs and AI chips | x86, chiplet design | Scaling AI and server share |
| Intel | Traditional server CPUs | x86, broad enterprise footprint | Trying to remain relevant in AI infrastructure |
What the Blackwell experience taught Nvidia
Nvidia is trying to avoid repeating the problems that affected its previous-generation Blackwell chips. Those chips reportedly ran into overheating issues when connected in the company’s customized server racks, forcing design changes and delaying shipments.
That history helps explain why Nvidia is emphasizing that Vera Rubin is more efficient, more integrated and easier to cool. The company appears determined to convince customers that its next major platform will avoid the kinds of headaches that can damage trust in a fast-moving market.
For a company that sells infrastructure at hyperscale, reliability is a product feature. Delays and thermal issues are not just engineering flaws; they can disrupt customer roadmaps and weaken confidence in future deliveries.
How big is the opportunity for Nvidia?
The opportunity is enormous because AI infrastructure spending is still expanding rapidly and remains highly concentrated among a relatively small number of buyers. The biggest hyperscalers and model developers are investing billions in capacity, and those customers want performance, efficiency and predictable supply.
If Nvidia can sell not just chips but integrated rack systems with CPUs, GPUs, software and cooling, it can potentially collect more revenue from each deployment and create stronger customer lock-in.
That business model also helps explain Nvidia’s persistent messaging around architecture and platform design. The company is not just trying to win a benchmark contest. It is trying to define the standard building block for the AI data center.
Key facts at a glance
- Vera Rubin is Nvidia’s next-generation AI system after Grace Blackwell.
- The NVL72 rack pairs 36 Vera CPUs with 72 Rubin GPUs.
- Nvidia says the platform delivers 10 times as many tokens per watt as Grace Blackwell.
- OpenAI already has one Vera Rubin rack in use, according to Nvidia.
- Early customers include Microsoft, OpenAI and Oracle.
Timeline of Nvidia’s Vera Rubin rollout
Nvidia has been steadily revealing more about Vera Rubin since its initial unveiling last year, but the company has been careful to present it as on schedule despite intense scrutiny.
| Date | Event | Why it matters |
|---|---|---|
| Spring 2025 | Nvidia first unveils Vera Rubin | Signals the next major platform after Grace Blackwell |
| Last week | Technical workshop in Santa Clara | Executives share new benchmarks and system details |
| This week | Public push ahead of AMD event | Sets up a competitive launch narrative |
| Second half of 2026 | Expected shipment window | Marks the next phase of Nvidia’s AI data center strategy |
What should customers and competitors watch next?
Customers should watch whether Nvidia can deliver Vera Rubin at scale without the thermal and integration issues that complicated earlier products. The company’s claims about speed, power efficiency and installation simplicity will matter most if they hold up in real deployments.
Competitors should watch how Nvidia packages CPUs and racks alongside GPUs. If buyers adopt the full system rather than individual chips, Nvidia could strengthen its position across the AI supply chain and make it harder for rivals to dislodge it.
For now, Vera Rubin is both a product launch and a strategic message. Nvidia is telling the market that the future of AI infrastructure belongs to companies that can design the whole machine, not just one piece of it.
And in the race to power the next generation of AI agents, Nvidia wants to own as much of that machine as possible.
Frequently asked questions
What is Nvidia’s Vera Rubin system?
Nvidia’s Vera Rubin system is the company’s next-generation AI hardware platform, combining Vera CPUs and Rubin GPUs in a rack-scale design. It is meant to power large AI data centers more efficiently than the current Grace Blackwell generation.
Why is Nvidia focusing on CPUs as well as GPUs?
Nvidia is focusing on CPUs because agentic AI systems need more orchestration, networking and data movement, tasks that often rely on central processors. The company wants to sell the full AI infrastructure stack, not just acceleration chips.
How does Vera Rubin compare with Grace Blackwell?
Nvidia says Vera Rubin is much more efficient than Grace Blackwell, with claims of 10 times as many tokens per watt and nearly three times the memory bandwidth. It also uses fewer cables and a more integrated liquid-cooled rack design.
Who is using Vera Rubin already?
OpenAI already has one Vera Rubin rack in use, according to Nvidia. The company also says Microsoft and Oracle are among the early customers expected to receive the system when it ships in the second half of the year.
Why is Nvidia announcing these details now?
Nvidia is releasing more information ahead of AMD’s annual product event to shape the market narrative and reinforce its lead in AI infrastructure. The timing also helps Nvidia highlight its own roadmap before a major competitor makes new announcements.









