Lisa Su's first rack-scale machine packs 72 GPUs and 31 terabytes of memory into a single liquid-cooled cabinet. For the first time in a decade, AMD is shipping into the same window as its rival rather than a year behind it.
SAN FRANCISCO — For most of the AI boom, the competitive landscape in data center compute has been described with a single number: somewhere between 70% and 85%. That is Nvidia's estimated share of the AI accelerator market by revenue, a grip so complete that "competition" has largely meant hyperscalers designing their own silicon rather than buying someone else's.
This week at the Moscone Center, AMD tried to change the arithmetic.
At its Advancing AI 2026 conference, chair and CEO Lisa Su walked a packed keynote hall through Helios — the company's first rack-scale AI system and, by her description, the industry's "highest-performance AI rack," built to train and run the most demanding frontier models at massive scale. It ships in the second half of this year, with a customer list that now includes Microsoft, Meta, OpenAI, Oracle, Anthropic and Tata Consultancy Services.
The pitch is no longer "a cheaper GPU." It is a full cabinet, sold as one unit, aimed squarely at the machine Nvidia sells at premium margins.
Helios is a double-wide, fully liquid-cooled rack built on the Open Rack Wide standard contributed to the Open Compute Project. Inside are 18 compute trays and six switches. Each tray pairs four Instinct MI455X accelerators with a single sixth-generation EPYC "Venice" CPU. That is 72 GPUs in total, every one of them sitting under a copper cold plate.
The headline specifications, per rack:
| Metric | AMD Helios | Nvidia Vera Rubin NVL72 |
|---|---|---|
| Accelerators | 72 × Instinct MI455X | 72 × Rubin |
| HBM4 per accelerator | up to 432 GB | up to 288 GB |
| HBM4 per rack | 31 TB | ~20.7 TB (derived) |
| FP4 inference | 2.9 exaflops | — |
| FP8 training | 1.4 exaflops | — |
| Scale-up bandwidth | 260 TB/s | 3.6 TB/s per GPU (NVLink 6) |
| Scale-out bandwidth | 43 TB/s | — |
| Interconnect | UALink over Ethernet + Ultra Ethernet | NVLink (proprietary) |
| Volume availability | 2H 2026 | 2H 2026 |
The MI455X itself is a monster of a chip: 320 billion transistors, eight compute dies on TSMC's 2nm process alongside I/O and fabric-and-cache dies on 3nm, twelve HBM4 stacks, and the distinction of being the largest part ever built on TSMC's CoWoS-L packaging. Peak throughput lands around 40 petaflops in MXFP4.
Memory is where AMD has planted its flag. A 50% capacity advantage per accelerator determines how many parameters stay resident on one GPU before inference has to fan out across neighbors, which is the bottleneck that most often wrecks serving economics on large models. On one leading open-weight model, AMD claims the MI455X delivers up to 34× the token throughput at high interactivity and up to 18× lower cost per token compared with its own previous-generation part.
Two caveats belong next to those figures. The two companies use different low-precision number formats that scale differently, so peak FLOPS comparisons are advertised ceilings rather than delivered performance. And Helios, per AMD's own documentation, is a reference design, a blueprint that OEM and ODM partners build branded systems from, not a SKU you order from AMD directly.
The most consequential architectural choice may be the one that doesn't show up in a benchmark. Inside the rack, all 72 GPUs share memory over UALink, a consortium fabric AMD runs on Ethernet. Between racks, it uses Ultra Ethernet Consortium specifications. Networking comes from the Pensando business AMD bought in 2022: the Vulcano 800 AI NIC and Salina 400 DPU, pushing 800 Gbps Ethernet against Nvidia's ConnectX and BlueField parts.
Nothing here is proprietary to AMD.
That is the entire argument: buyers assembling gigawatt-scale fleets have watched a single vendor's interconnect become the thing that locks them in, and AMD is selling the exit.
The customer roster is doing heavy lifting for the narrative. Microsoft confirmed it will deploy Helios across Azure, joining Meta, OpenAI and Oracle. AMD and Anthropic announced a strategic partnership to deploy up to two gigawatts of GPUs on the platform. OpenAI and Meta have separately committed to a combined 12 gigawatts of AMD accelerator capacity. AMD says eight of the world's ten largest AI companies now run workloads on Instinct silicon.
Pricing tells its own story. Analysts estimate Helios racks will land between $5 million and $5.5 million, against roughly $3.5 million to $4 million for the competing Nvidia system, a premium of about 40%. AMD is not undercutting anyone. It argues the higher peak performance yields roughly a 30% advantage in performance per dollar, and, notably, Microsoft bought anyway.
The financial gap those racks are meant to close remains enormous. Nvidia's most recent quarter brought in $81.6 billion in total revenue, $75.2 billion of it from data center, against AMD's $5.8 billion data center quarter, itself up 57% year over year. Put differently: Nvidia's data center business alone is roughly thirteen times the size of AMD's entire data center franchise.
But AMD's trajectory is steepening. Total revenue rose 38% to $10.3 billion in the most recent reported quarter, free cash flow hit a record $2.6 billion, and guidance for the following quarter sits near $11.2 billion, about 46% growth. Management expects tens of billions of dollars in data center AI revenue starting in 2027, most of it from Helios, and now sizes the accelerator opportunity at $1.4 trillion. The company's stated goal is double-digit market share within three to five years, up from an estimated 5% to 8% today.
Wall Street has been repricing accordingly. Shares traded near $553 during the two-day event, more than double where they started 2026, and price targets were raised into the $600s and $700s across multiple firms. Roughly 82% of covering analysts rate the stock a buy, with no sell ratings.
Three things, mainly.
Timing. At least one prominent research shop has argued that volume production slips to the second quarter of 2027. AMD flatly denies it, guiding instead to initial volume in Q3 2026 with a larger ramp in Q4 and into early next year. Nothing settles this except racks arriving on loading docks.
Software. ROCm has improved substantially, but the out-of-box experience still tends to require real kernel engineering to close the utilization gap against a CUDA ecosystem with a fifteen-year head start. Buying a faster rack is easy. Rewriting a serving stack is not.
The third competitor. Custom silicon from the cloud giants is growing far faster than merchant GPUs, roughly 45% year over year against 16%, and now accounts for a meaningful slice of accelerator server shipments. AMD is fighting for second place in a market where the largest customers are increasingly their own suppliers.
Here is what separates this launch from AMD's previous attempts. In past generations, AMD would post a memory or FLOPS advantage and then arrive twelve months after the competing part had already been designed into everyone's data centers. The advantage was real and irrelevant.
This time both platforms reach volume in the same half of the same year. AMD has more memory per rack, open interconnects, a hyperscaler anchor tenant, gigawatt-scale commitments from the largest model developers on earth, and a price that assumes it doesn't need to discount.

Comments