Microsoft Bets on AMD's Helios for Azure AI

The compute hunger keeps growing

Every few months, the AI world produces a new model that everyone suddenly wants to run, tweak, and serve to their own users. The latest driver is a wave of freely available, top-tier open models that anyone can fine-tune. That access is great for builders, but it comes with a catch: you need serious hardware to run the stuff. And the appetite for that hardware, measured in FLOPS (floating-point operations per second, basically how fast a chip does math), shows no sign of slowing down.

Into that backdrop steps a new deal. Microsoft and AMD announced that Microsoft will add AMD's Helios rack-scale AI accelerator in volume, both inside its own data centers and for customers using Azure, its cloud platform.

What Helios actually is

Helios is not a single chip. It is a whole rack of them working as one system, which is why it gets the label "rack-scale." It bundles 72 of AMD's next-generation Instinct MI455X GPUs together, with a combined 31.1TB of HBM4 memory. HBM is the ultra-fast memory that sits right next to AI chips so they are not left waiting for data.

The performance numbers are large. AMD claims up to 1.4 exaFLOPS of FP8 compute and 2.9 exaFLOPS of FP4. FP8 and FP4 are compact number formats that trade a little precision for a lot of speed, which is exactly the trade AI workloads like. An exaFLOP is a quintillion operations per second, so these are big figures. AMD is also targeting 260 TB/s of internal, or "scale-up," bandwidth within the rack, and 43 TB/s of "scale-out" bandwidth for linking racks together.

The obvious target here is Nvidia. Helios is built to compete with Nvidia's Vera Rubin NVL72 system, expected later this year. AMD says its internal bandwidth is on par with Nvidia's and its scale-out figure is roughly double, though that last number relies on a networking approach called UALink over Ethernet whose real-world performance is not yet proven.

Why the deal matters

Nvidia has dominated AI data center hardware for years. For AMD, landing a hyperscaler like Microsoft is a meaningful step toward chipping away at that lead. It follows other large AMD partnerships over the past year with OpenAI and Meta, deals measured in gigawatts of compute and potentially hundreds of billions of dollars.

For Azure customers, the practical upshot is choice. AI labs get another option for training and running models, and enterprises can tap this compute through Microsoft Foundry, the company's managed platform for building AI applications. More competition at the hardware level tends to help buyers, whether through pricing, availability, or simply not being locked into one supplier.

One honest caveat: neither company said how big the deployment actually is. There were no figures in watts or dollars, so "at scale" and "in volume" are doing a lot of work here. The performance claims also come from AMD, not from independent testing.

It is not just GPUs

The announcement went beyond Helios. Azure will add two new virtual machine types built on AMD's upcoming sixth-generation Epyc Venice server processors. A virtual machine is a slice of cloud hardware you rent as if it were your own computer. The HDv2 series targets "agentic AI and data pipelines," meaning AI systems that take actions on their own, while the HXv2 series is aimed at chip design work. Microsoft also plans to fold AMD's Pensando DPUs, specialized chips that handle networking and storage tasks, deeper into its Azure Boost service.

What is next

The real test comes when Helios ships later this year and runs head to head against Nvidia's newest systems in live data centers. That is when AMD's paper specifications, especially the unproven scale-out networking, meet reality. AMD is expected to share more at its Advancing AI event this week. For now, the takeaway is simpler: the AI hardware race just got a little less one-sided, and the biggest cloud providers are increasingly happy to hedge their bets.