Tue 21 Jul 2026 / 16:26 ET
Kernel
Internet 4 min read

Nvidia pitches Vera Rubin as its bid for the whole AI rack

Nvidia says its next AI system pairs Rubin GPUs with Vera CPUs as demand grows for processors that coordinate agentic workloads.

June Castellano

By June Castellano / Platforms & Power Reporter

Nvidia pitches Vera Rubin as its bid for the whole AI rack
img: WIRED

Nvidia is using its Vera Rubin rollout to make a blunt pitch to AI data center buyers: stop thinking of the company as a GPU vendor and start treating it as the supplier for the whole rack.

At a technical briefing at Nvidia’s Santa Clara headquarters last week, company executives laid out new performance claims for Vera Rubin, its next CPU and GPU system for AI workloads, according to WIRED. The timing was not subtle. AMD is holding its annual product event in San Francisco on Thursday, where it is expected to promote its own AI and data center hardware.

The shift matters because AI infrastructure is no longer just a contest over the fastest accelerator. GPUs still do the heavy lifting for training and running large models. But more complex agent-style systems also need CPUs to manage data movement, networking, and software orchestration. Nvidia wants that part of the bill of materials too, a perfectly ordinary ambition dressed up as architecture strategy.

What Nvidia is selling

Vera Rubin follows Nvidia’s Grace Blackwell system. The design pairs one Vera CPU with two Rubin GPUs. In the Vera Rubin NVL72 rack-scale system, Nvidia says 36 Vera CPUs sit alongside 72 Rubin GPUs in a liquid-cooled platform.

Nvidia is also offering the Vera CPU by itself. Reuters reported that Nvidia has told customers in China that standalone Vera CPUs could be available as early as August.

During a data center lab tour, Nvidia executives said OpenAI already has one Vera Rubin rack in use, according to WIRED. Nvidia CEO Jensen Huang did not attend the Santa Clara briefing. He was in Japan announcing AI robotics partnerships with Japanese companies, Reuters reported.

Ian Buck, Nvidia’s vice president of accelerated computing and a key figure behind CUDA, led the briefing. Buck told reporters the company is building a roadmap that includes new CPU architectures as well as GPUs, adding that continued architectural work is a survival requirement in Silicon Valley.

The benchmark claims need context

Nvidia claims Vera Rubin NVL72 can process 10 times more tokens per watt than Grace Blackwell. The company also says the Vera CPU beats AMD and Intel CPUs on agentic AI workloads. WIRED noted that the comparisons appeared to use somewhat older rival processors, which is the sort of footnote that tends to do real work in vendor benchmarks.

Nvidia says Vera Rubin’s localized memory subsystems provide nearly three times the memory bandwidth of Blackwell. That could appeal to hyperscalers trying to buy around constrained supplies of high-bandwidth memory.

The company is also promoting the rack design as easier to install and service. Buck and Andrew Bell, Nvidia’s senior vice president of hardware engineering, said reduced cabling could cut rack installation time from hours to minutes. Nvidia says the system is fully liquid-cooled, which can reduce the energy needed for cooling compared with air-based approaches.

Huang has said Vera Rubin is moving toward full production and will ship in the second half of this year. Nvidia has named Microsoft, OpenAI, and Oracle among early customers. The company has reason to be prickly about schedules: its prior Blackwell systems reportedly suffered overheating issues when chips were connected in customized server racks, requiring design changes and delaying shipments.

AMD is waiting on the other side

AMD has been gaining share in data center CPUs over the past two years and remains a major x86 supplier. On Sunday, AMD disclosed more about Helios, an AI rack intended to compete with Nvidia’s systems, CNBC reported.

The two companies are chasing long, high-volume chip deals with hyperscalers and AI labs including Meta, Amazon, OpenAI, Anthropic, and SpaceXAI, according to WIRED.

Nvidia’s CPU approach differs from AMD’s in two major ways described by its executives. Vera uses Arm rather than x86, and Nvidia says it uses a monolithic chip instead of the chiplet designs common in modern processors. Hannah Coutand, who runs product marketing for Nvidia DGX Cloud, argued that chiplets impose costs on memory bandwidth and data movement, while Vera’s single-die design lets data move faster across one integrated circuit.

This story draws on original reporting from WIRED.

More Internet/

view all ↗