Microsoft will put AMD’s next-generation Helios AI rack systems into Azure data centers “at scale,” giving AMD another shot at the part of the cloud AI market that Nvidia has largely owned. Microsoft and AMD said the deployment will support Microsoft’s own frontier-model work, Azure AI infrastructure customers, and AI services sold through the cloud platform.
The companies did not say how many racks Microsoft will buy, how much power the deployment will consume, or what the deal is worth. That omission matters. In AI infrastructure, “at scale” can mean very different things depending on whether a company is talking about a showcase cluster, a regional cloud buildout, or enough hardware to dent Nvidia’s lead.
What Microsoft and AMD did say is more concrete: Azure customers, including AI labs, will be able to use the new AMD systems for training models and serving inference. Microsoft also plans to use the hardware behind managed AI compute for enterprise customers deploying workloads through Microsoft Foundry.
What Helios is
AMD’s Helios is a rack-scale AI system, meaning AMD is selling the rack as the unit of compute rather than treating each accelerator server as a separate box. The design ties together 72 AMD Instinct MI455X GPUs in one system. AMD says a full Helios rack provides 31.1TB of HBM4 memory across those GPUs.
For lower-precision AI math, AMD says Helios can reach up to 1.4 exaFLOPS of FP8 compute and 2.9 exaFLOPS of FP4 compute using OCP AI data types. Those formats trade numerical precision for speed and memory efficiency, which is often acceptable for parts of training and inference when the model and software stack are built for it.
AMD is aiming Helios at the same class of deployment as Nvidia’s Vera Rubin NVL72 system, which is expected later this year. AMD says Helios targets 260TB/s of scale-up bandwidth inside the rack, matching the figure cited for Nvidia’s rack-scale system. For scale-out links between racks, AMD says Helios targets 43TB/s using UALink over Ethernet, roughly twice the figure cited for Vera Rubin. The useful caveat is the boring one: UALink over Ethernet performance in production clusters still has to prove itself outside vendor slides.
Azure also gets new AMD CPU instances
Microsoft and AMD also said Azure will add two virtual machine series based on AMD’s upcoming sixth-generation Epyc “Venice” CPUs. Microsoft described the HDv2 series as aimed at “agentic AI and data pipelines,” while the HXv2 series is intended for semiconductor design workloads.
Microsoft also plans to use its existing AMD Pensando data processing units in Azure Boost, the company’s system for moving networking and storage work off the main CPU. In plain terms, the DPUs handle infrastructure chores that would otherwise eat into host processor time.
The announcement adds Microsoft to a list of major AMD AI infrastructure customers at a time when AMD is trying to take more data center GPU share from Nvidia. AMD has also struck large AI compute partnerships with OpenAI and Meta in the past year, according to company announcements. Microsoft’s commitment gives AMD another high-profile cloud deployment, though the missing deployment size keeps the actual scale fuzzy for now.
This story draws on original reporting from Tom's Hardware.