AMD’s Helios System Takes Direct Aim at Nvidia’s AI Dominance With Rack-Scale Power and Surprising Alliances

AMD unveiled Helios, a rack-scale AI system with 72 MI455X GPUs delivering 2.9 exaFLOPS FP4 and 31 TB HBM4. It challenges Nvidia's Vera Rubin while partnering with Cerebras for low-latency inference. Microsoft, OpenAI, Meta and Anthropic commit to deployments starting late 2026. The platform promises 30% better tokens per dollar and up to 5x efficiency gains in mixed workloads. AMD forecasts a $1.4T accelerator market by 2030.
AMD’s Helios System Takes Direct Aim at Nvidia’s AI Dominance With Rack-Scale Power and Surprising Alliances
Written by Ava Callegari

AMD just dropped its most ambitious bid yet to crack Nvidia’s stranglehold on AI infrastructure. The company unveiled Helios, a full rack-scale system built around the new Instinct MI455X accelerators. And the timing couldn’t be more pointed. Nvidia’s Vera Rubin platform has set the bar high. But early customers already line up for AMD’s answer.

Details emerged at the Advancing AI event in late July 2026. Helios packs 72 MI455X GPUs, each with 432 GB of HBM4 memory. That’s 31 TB total across the rack. Bandwidth hits 1.4 PB per second aggregate. The system delivers up to 2.9 exaFLOPS of FP4 performance. Or 1.4 exaFLOPS at FP8. These numbers come straight from AMD’s own disclosures.

But raw specs tell only part of the story. Helios integrates AMD’s sixth-generation EPYC Venice CPUs, with up to 256 cores each. It adds Pensando networking chips acquired years ago. The entire rack weighs around 5,000 to 7,000 pounds. It draws 225 to 245 kilowatts. Liquid cooling keeps it in check. And it follows Meta’s Open Compute Project standards for easier adoption.

Yet the real shift lies in AMD’s willingness to partner outside its walls.

The standout move pairs Helios with Cerebras Systems’ wafer-scale engines. This disaggregated setup splits inference workloads. Helios handles high-throughput prompt processing. Cerebras delivers ultra-low-latency token generation. Together they promise up to 5 times higher tokens per second per watt on models like Kimi 2.6. Modeling from AMD and Cerebras labs in July 2026 backs the claim. The joint solution reaches customers through Cerebras Cloud in the second half of 2026. Cerebras will also deploy Helios racks in its own data centers.

“AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach,” said AMD CEO Lisa Su in the Cerebras press release. “AMD Helios delivers leadership performance and scale for the broadest range of inference workloads. Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI.”

Andrew Feldman, Cerebras CEO, struck a similar tone. “The demand for ultra-fast inference is growing at an unprecedented pace. Cerebras delivers the world’s fastest, ultra-low-latency inference. Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers,” he added in the same release.

This alliance stands out. It signals AMD no longer tries to match Nvidia single-handedly across every workload. Instead it mixes architectures where each shines. Recent X posts from industry watchers highlight the logic. One noted how Helios manages prefill while Cerebras accelerates decode. The combination targets real-time agents, coding copilots and robotics where response time decides value.

Customer momentum builds fast. Microsoft committed to deploy Helios in Azure data centers. Satya Nadella spelled it out. “We are expanding the Azure infrastructure portfolio with AMD Helios to give customers the performance, scale and choice they need to build and run the next generation of AI applications,” the Microsoft CEO said, according to CNBC coverage.

OpenAI plans massive-scale deployment starting late 2026 into 2027. It already eyes the follow-on MI500 series. Anthropic committed up to 2 gigawatts of AMD GPUs, with initial Helios racks arriving in the first half of 2027. Meta signed on for up to 6 gigawatts over time, beginning with 1 gigawatt this year. Oracle, Tata Consultancy Services and others joined the list. Eight of the top 10 AI developers already run workloads on AMD Instinct GPUs.

Forrest Norrod, AMD’s data center chief, didn’t mince words. He called Helios “our baby.” And he stressed the focus. “We’re very focused on providing the best total cost of ownership, the lowest cost per token, all in,” Norrod told CNBC. “And our customers are telling us that we’re achieving that.”

Analysts peg Helios rack cost between $5 million and $5.5 million. Nvidia’s Vera Rubin equivalent runs an estimated $3.5 million to $4 million. The premium buys 50 percent more HBM4 capacity. It delivers higher scale-out bandwidth at 43 TB/s versus 28.8 TB/s. In moderate interaction scenarios with models like Kimi2, Helios shows 15 percent higher GPU throughput. AMD claims up to 30 percent more tokens per dollar overall. And 36 times the performance of its prior generation.

Yet challenges remain. Nvidia’s CUDA software still commands broader developer loyalty. AMD pushes ROCm as an open alternative. It now supports PyTorch, JAX, ONNX and vLLM with day-zero model compatibility. But ecosystem inertia favors the leader. Neil Shah, analyst at Counterpoint Research, captured the tension. AMD’s chips sit “on par” with Nvidia’s. “But the secret sauce is in the software and optimization.”

Daniel Newman of the Futurum Group offered a bullish take. He sees AMD potentially capturing 20 to 25 percent market share. That slice represents hundreds of billions in revenue. AMD itself forecasts the AI accelerator market alone will hit $1.4 trillion by 2030. Total computing spend could reach $2 trillion. No single vendor solves every problem, Su has noted repeatedly.

The MI455X itself pushes boundaries. Built on a mix of 2nm and 3nm processes from TSMC, the chip contains 320 billion transistors. It offers 40 petaFLOPS of FP4 compute. Memory bandwidth reaches 23.3 TB/s. That’s more than double the previous generation. Each GPU holds 432 GB of HBM4, a 50 percent jump over Nvidia’s Rubin equivalents in capacity though peak FLOPS trail slightly in some precisions.

Venice EPYC CPUs complement the GPUs. These Zen 6 chips enter volume production on TSMC’s 2nm process first among high-performance parts. They deliver over 70 percent better performance per efficiency. Single-core gains exceed 20 percent. Thread density rises more than 30 percent. In agentic AI benchmarks they post 2.2 times the throughput of comparable Nvidia offerings in some tests.

Shipping begins late 2026. Volume deployments ramp in the second half. AMD books data center revenue in the tens of billions starting 2027, with Helios driving most of it. The company raised its long-term accelerator market view on the strength of these systems.

Recent coverage reinforces the stakes. Wccftech reported on the 72-GPU configuration and open standards push. It highlighted how Helios exceeds Vera Rubin in memory density and scale-out bandwidth while embracing UALink over Ethernet for initial interconnects. Ethernet may introduce some latency trade-offs compared with Nvidia’s NVLink, but it eases multi-vendor integration.

The Next Web framed the announcement as AMD advancing on multiple fronts at once. It noted OpenAI’s Sachin Katti praising plans to deploy Helios at massive scale. Anthropic’s quick adaptation of Claude to the racks over a single weekend also drew attention. Shares dipped 4 percent after the event despite 145 percent gains year to date. Investors weighed the competition ahead.

AMD’s road map extends further. The MI500 series arrives next year with HBM4e and new copper-optical interconnects. Inference throughput could exceed 2,000 times that of the MI300X from four years prior. MI600 follows in 2028. Helios 500 and 600 variants will track those GPU generations. On the CPU side, Florence and Ravenna successors target 2028 and 2030.

The bet feels clear. AMD refuses to fight Nvidia solely on proprietary software lock-in. It offers open standards, heterogeneous computing and aggressive memory scaling. It courts the biggest AI labs with tangible cost-per-token wins. Whether that erodes Nvidia’s lead depends on execution. Early deployments will test software maturity and real-world efficiency claims.

But the message lands. The AI infrastructure race just gained a formidable second contender. One that ships systems this year. One that partners with wafer-scale specialists. And one that counts Microsoft, OpenAI and Meta among its first buyers. The sun god Helios pulls no single horse. It harnesses four in-house strengths plus external allies. Nvidia now faces its most credible rack-scale rival in years.

Subscribe for Updates

EmergingTechUpdate Newsletter

The latest news and trends in emerging technologies.

By signing up for our newsletter you agree to receive content related to ientry.com / webpronews.com and our affiliate partners. For additional information refer to our terms of service.

Notice an error?

Help us improve our content by reporting any issues you find.

Get the WebProNews newsletter delivered to your inbox

Get the free daily newsletter read by decision makers

Subscribe
Advertise with Us

Ready to get started?

Get our media kit

Advertise with Us