AMD and Cerebras are combining their hardware into a single AI inference system that promises to undercut Nvidia on latency and cost, the companies said Thursday.
AMD and Cerebras are combining their hardware into a single AI inference system that promises to undercut Nvidia on latency and cost, the companies said Thursday.

AMD and Cerebras are combining their hardware into a single AI inference system that promises to undercut Nvidia on latency and cost, the companies said Thursday.
"This is the first time anyone has disaggregated inference this way — using Cerebras for the ultra-low-latency decode and AMD for the high-throughput prefill," Andrew Feldman, chief executive officer of Cerebras, said in an interview.
The combined solution will pair AMD Helios racks — each containing 72 Instinct MI455X GPUs and 18 sixth-generation Epyc "Venice" CPUs — with Cerebras wafer-scale engines optimized for rapid token generation. AMD claims Helios delivers up to 30% more inference tokens per dollar than the leading competitive solution, a direct reference to Nvidia's Vera Rubin NVL72 platform now in full production and shipping to customers including OpenAI, CoreWeave and Microsoft. The system will ship from Cerebras-owned data centers starting in the fourth quarter, Feldman said.
The partnership gives AMD a differentiated offering in the fastest-growing segment of AI infrastructure — inference — where Nvidia's dominance is less entrenched than in training. AMD expects the total addressable market for AI accelerators to exceed $1 trillion by 2030 as workloads shift from model training to real-time reasoning and agentic AI.
How the Architecture Works
The system splits inference into two stages. AMD Helios handles the prefill phase — processing the user's query against the model's context — while Cerebras wafer-scale engines generate tokens at speeds measured in milliseconds. The approach targets customers running applications where response time is critical, such as conversational agents, coding assistants and real-time fraud detection. Nvidia has been assembling a similar capability through its recent acquisition of Groq, a startup specializing in low-latency inference chips.
Partners Already Lining Up
OpenAI expects to bring Helios online beginning in the fourth quarter of 2026, with deployments accelerating through 2027, Sachin Katti, a senior executive at OpenAI, said on stage at AMD's Advancing AI event. Anthropic, which received a commitment of up to $5 billion from AMD on Wednesday, plans to deploy 2 gigawatts of MI450-series GPUs in Helios racks to run its Claude models, with the first gigawatt shipping in the first half of 2027. Meta is also validating Helios racks for gigawatt-scale deployments.
Investor Impact
AMD shares fell about 4% during the presentation despite having more than doubled this year, suggesting Wall Street had priced in ambitious product launches. Nvidia's Vera Rubin platform is already generating revenue, while AMD's Helios is just entering production. The Cerebras partnership gives AMD a time-to-market advantage in low-latency inference, but the revenue contribution will depend on how quickly customers adopt the disaggregated architecture. AMD trades at roughly 22 times forward earnings, a discount to Nvidia's 35 times, reflecting the gap in execution credibility that the company is trying to close with partnerships like this one.
This article is for informational purposes only and does not constitute investment advice.