AMD is making its boldest move yet to loosen NVIDIA’s grip on the AI chip market. At its Advancing AI event in San Francisco on July 23, chief executive Lisa Su launched Helios, the company’s first rack-scale AI system. AMD says it is already in production, and Reuters reports shipments should begin near the end of the third quarter.
“The next phase of AI will span frontier models, agents and physical AI, creating new opportunities to bring intelligence everywhere,” said Dr. Lisa Su, chair and CEO, AMD. “Realizing that potential will take the entire industry working together. AMD is partnering across the ecosystem to deliver leadership compute and open platforms that give customers the performance, flexibility and choice to scale AI from the data center to the edge.”
Each Helios rack combines 72 Instinct MI455X GPUs with 18 sixth-generation EPYC “Venice” CPUs, tied together by AMD’s Pensando networking and its ROCm software stack. That design matters, since it moves AMD’s fight with NVIDIA beyond single chips toward a complete, integrated rack where processors, memory, networking, and software are engineered as one.
The whole system targets inference computing. Inference is the data crunching that happens every time someone queries a chatbot like ChatGPT, and it is the fastest-growing slice of AI demand. AMD claims Helios can deliver up to 30% more inference tokens per dollar than a leading rival system. That figure is a vendor estimate based on a projected workload, though, so it still needs independent validation.
Helios did not arrive alone, either. AMD also unveiled sixth-generation EPYC CPUs, the Instinct MI400 GPU series, Ryzen AI Embedded X100 processors, and a Kria robotics development platform. Together they signal a full-stack push across data centers, edge devices, and robotics rather than a single flagship launch.
The customer names give the platform real weight, and several partners detailed how they will deploy AMD infrastructure at scale. Anthropic, the maker of Claude and the company behind this article’s own AI, expanded on Wednesday’s deal to deploy up to two gigawatts of MI455X GPUs in Helios racks. The two are also starting a multiyear engineering collaboration to use Claude itself to optimize workloads for AMD Instinct GPUs and speed up ROCm software development, and AMD will adopt Claude across its own engineering teams. Anthropic co-founder Tom Brown then shared a remarkable detail, saying Claude had set up AMD’s AI servers entirely on its own over a single weekend.
“Anyone, human or AI, can now build real models on your platform,” he said, hinting at a future where AI configures the very hardware it runs on.
OpenAI is going deeper on the software side. It is pairing its Triton framework with AMD’s ROCm to optimize GPT-class workloads on MI455X GPUs and Helios racks. OpenAI expects to bring Helios online from the fourth quarter of 2026, with deployments accelerating through 2027, building on an October deal giving it the option to buy up to roughly 10% of AMD.
The other two names carry serious weight too. Meta is co-designing with AMD for gigawatt-scale deployments, and it has begun validating both sixth-generation EPYC platforms and Helios racks in its labs ahead of scaling up. Cerebras, meanwhile, is combining its ultra-low-latency compute with Helios’s high-throughput infrastructure to improve the economics of ultra-low-latency inference serving.
Su was blunt about AMD’s ambitions beyond second place. She said the company is making another major leap and expects leadership in the scale-up compute domain. She projected the total computing market will reach $2 trillion by 2030, with $1.4 trillion of that from AI-accelerating chips.
For all the ambition, the market stayed skeptical on the day. AMD shares fell about 2% amid a broad semiconductor selloff, while Cerebras rose around 5%. The bigger test comes next, since the real proof will be reproducible benchmarks, software maturity across tools like PyTorch and vLLM, and delivery at scale rather than launch-day claims.
