TradingKey - AMD ( AMD )'s plan to challenge Nvidia ( NVDA )'s dominance in AI chips is taking a crucial step forward.
On July 20 local time, Microsoft ( MSFT) announced that it will deploy AMD's next-generation rack-scale AI system Helios in its Azure data centers to provide computing power support for frontier large model inference and Azure AI services. This not only means Microsoft is further expanding its partnership with AMD, but also allows Helios, following Meta ( META ), OpenAI, and Oracle ( ORCL ), to secure another top-tier global cloud service provider as a client.
For AMD, the significance of this partnership goes far beyond just selling an AI server system.
Microsoft stated that Helios will become a key component of Azure AI infrastructure, primarily handling large model inference tasks and providing enterprise customers with the computing power required for next-generation AI applications.
Meanwhile, Azure will also launch two new compute instances based on AMD's sixth-generation EPYC Venice processors, targeting scenarios such as Agent AI, data processing, and semiconductor design, while also deploying AMD Pensando DPUs in select Azure services to further enhance network-layer capabilities.
This means AMD is no longer providing Microsoft with just GPUs, but a comprehensive AI infrastructure solution spanning GPUs, CPUs, networking chips, and software platforms.
Microsoft CEO Satya Nadella stated that the company hopes to leverage AMD Helios to further expand Azure's computing options, providing customers with higher-performance, larger-scale, and more flexible AI infrastructure.
In fact, Microsoft and AMD have a long history of collaboration. From Surface PCs and Xbox consoles to pioneering the adoption of the MI300X GPU in 2023, and now deploying Helios, their partnership has steadily extended into AI data centers. As Microsoft continues to ramp up self-developed large models and expand the scale of its Azure AI business, its need for a diversified supply of computing power has become increasingly urgent.
For Microsoft, introducing AMD not only eases the tight supply of AI GPUs but also helps reduce dependence on a single supplier.
If the MI300X was still a competition at the GPU product level, then Helios signifies AMD officially launching a challenge against Nvidia's integrated datacenter system solutions.
As AMD's first rack-scale AI system, Helios integrates Instinct GPUs, EPYC Venice CPUs, Pensando networking chips, and the ROCm software platform, achieving integrated synergy across the four core capabilities of GPU, CPU, networking, and software.
Forrest Norrod, head of AMD's datacenter business, stated that the company hopes to achieve lower total cost of ownership (TCO) and lower per-token inference costs through system-level optimization, rather than just pursuing single-GPU performance.
Previously, AMD CEO Lisa Su also stated that Helios holds a clear advantage over Nvidia's Grace Blackwell and next-generation Vera Rubin systems in terms of inference performance, memory capacity, and bandwidth.
Notably, Helios has not adopted a low-price competition strategy.
According to estimates by Futurum Group, a single Helios system is priced at approximately $5 million to $5.5 million, which is higher than Nvidia's Vera Rubin at about $3.5 million to $4 million. This means AMD is attempting to compete for market share through overall performance, system efficiency, and long-term operating costs, rather than a price advantage.
It is widely believed in the industry that AI infrastructure competition is escalating from individual GPU products to a showdown between integrated AI systems, and Helios is precisely AMD's first trump card in this arena.
Currently, companies including Meta, OpenAI, Oracle, and Indian IT services giant TCS have all announced the deployment of Helios. Among them, Meta plans to cumulatively adopt up to 6GW of AMD GPU compute capacity in the future, with the first 1GW to be deployed first on the Helios platform; OpenAI and Oracle will also begin deploying related systems this year.
AMD stated that eight of the world's top ten AI companies currently run workloads on its Instinct GPU platform, signaling that the company is rapidly penetrating the AI training and inference markets.
However, looking at the overall market, the gap between the two remains significant.
Currently, Nvidia still holds over 95% of the global data center GPU market share, while AMD accounts for only about 4.5%.
Daniel Newman, CEO of Futurum Group, believes that if Helios can successfully complete large-scale deployment and gain customer recognition, AMD's data center GPU market share is expected to rise to 20% to 25% in the coming years, representing tens of billions of dollars in new revenue potential.
AMD also expects that starting in 2027, its data center AI business revenue will reach tens of billions of dollars, with the vast majority of the growth coming from the Helios platform.
Although Helios has made breakthroughs at the hardware level, the market generally believes that AMD still faces one last hurdle before truly challenging Nvidia—the software ecosystem.
Over the past two decades, Nvidia has built a complete software development system around CUDA, including TensorRT, Dynamo, and development tools continuously optimized for inference scenarios, allowing developers to quickly complete model deployment and performance tuning.
In contrast, AMD has continuously invested in building the ROCm ecosystem in recent years and has bolstered its system integration capabilities through acquisitions of ZT Systems, Pensando, and several software companies, though there is still room for improvement in overall software maturity.
Neil Shah, an analyst at Counterpoint Research, pointed out that while AMD is already close to Nvidia in terms of GPU and CPU performance, software compatibility, development tools, and ecosystem completeness remain critical factors in determining enterprises' final purchasing decisions.
For large cloud providers, purchasing AI infrastructure is not just about chip performance; they place greater emphasis on overall deployment efficiency, software support, development costs, and long-term maintenance capabilities.
Therefore, whether Helios can continue to secure more orders in the future will largely depend on the actual deployment outcomes of the initial customers and the development speed of the ROCm ecosystem.