Nvidia is talking about CPUs with unusual enthusiasm for a company whose money machine is still the GPU. Ian Buck, Nvidia’s vice president of hyperscale and high-performance computing and the inventor of CUDA, said the company has shipped standalone Grace CPU servers at a scale “in the hundreds of thousands,” according to Tom’s Hardware.
That matters because Nvidia is trying to sell more than accelerator cards into AI data centers. In May, the company said it had shipped more than 2.5 million Grace CPUs overall. In February, Nvidia also announced that Meta would put standalone Grace systems into production, with Vera systems to follow.
The shift in the pitch follows a change in the work. Tom’s Hardware reported that agentic AI workloads are altering server designs that once paired many GPUs with a single CPU. In some cases, the balance is moving closer to one CPU per GPU. That gives Nvidia a reason to argue that its own CPUs belong in the same buying conversation as AMD Epyc and Intel Xeon parts.
Grace opened the CPU door for Nvidia
Grace is Nvidia’s first real data center CPU beachhead. The chip uses 72 Arm Neoverse V2 cores. Its differentiator, according to Nvidia, is the company’s Scalable Coherency Fabric, which is meant to let cores, cache and memory controllers communicate with high bandwidth and low contention.
Buck said Grace systems were not aimed at cheap cloud hosting. “They weren’t running a web server,” he said, according to Tom’s Hardware. He described the deployments as backend systems for data-heavy jobs, including data processing.
Vera is the more aggressive follow-up. Nvidia says Vera uses an updated coherency fabric and its first custom CPU core design, called Olympus. The chip has 88 cores on one monolithic die, rather than the chiplet approach AMD and Intel moved toward years ago for their largest server processors.
That design choice is the technical bet. Chiplets help vendors pack more cores into a package, but they add latency and coherency headaches. Nvidia is spending die area on the on-chip fabric instead. Buck said Vera has 3.4 TB/s of core-to-core bandwidth inside the CPU so each core can reach cache and memory controllers without collisions. Tom’s Hardware clarified that Vera’s aggregate memory bandwidth is up to 1.2 TB/s through LPDDR5X, or 14 GB/s per core.
The trade-off is the old server world
Nvidia is not claiming one CPU will fit every data center job. Buck said at Nvidia’s GTC event in March that “the world is not going to be served by one SKU of CPU,” according to Tom’s Hardware. He also acknowledged that spending silicon on Vera’s fabric limits the core count, saying that is one reason the chip has 88 cores rather than 128.
The open question is how much of the market wants that trade. Tom’s Hardware noted that many data center workloads remain legacy jobs built around existing hyperscaler infrastructure. Nvidia’s answer is that agentic AI expands the CPU opportunity and changes the requirements.
The company puts the CPU total addressable market at $200 billion. Tom’s Hardware reported that other industry estimates have been closer to $120 billion by 2030, with some recent figures as high as $170 billion. Morgan Stanley estimated in April that agents could add as much as $60 billion to the data center CPU market.
Vera is now in full production alongside Nvidia’s next AI infrastructure generation, including Rubin GPUs, ConnectX-9 network interface cards and SpectrumX Ethernet switches. Nvidia says a Vera Rubin NVL72 rack contains about 1.3 million components and that more than 300 partners are involved globally in building the systems. Tom’s Hardware reported seeing a Vera Rubin NVL72 rack running OpenAI workloads during a visit to Nvidia’s headquarters.
This story draws on original reporting from Tom's Hardware.