AMD And Cerebras Team Up On Disaggregated AI Inference With Helios And WSE

AMD and Cerebras Systems unveiled a joint disaggregated inference platform at Advancing AI 2026, pairing AMD Helios rackscale GPU systems with the Cerebras Wafer-Scale Engine to promise up to 5x higher tokens per second per watt.

AMD And Cerebras Team Up On Disaggregated AI Inference With Helios And WSE

AMD (NASDAQ: AMD) and Cerebras Systems (NASDAQ: CBRS) announced a technical partnership on July 23, 2026 to deliver a disaggregated AI inference platform that combines AMD Helios rackscale solutions with the Cerebras Wafer-Scale Engine. The joint solution, unveiled at AMD's Advancing AI 2026 event in San Francisco, is designed to serve latency-sensitive inference workloads such as coding copilots, real-time voice agents and agentic tool use without giving up the throughput and scale of a modern GPU rack.

How Disaggregation Works

Modern LLM inference has two very different stages. Prompt processing crunches long contexts and reasoning traces and benefits from GPU throughput; decoding generates the actual tokens one at a time and is memory-bandwidth bound. AMD and Cerebras will run those stages on different silicon in a single workflow. AMD Helios racks, built on AMD Instinct MI400 series accelerators, HBM4 memory and Pensando networking, provide the high-throughput prompt engine. The Cerebras Wafer-Scale Engine, which packs 900,000 AI cores and 44 GB of on-die SRAM into a single 46,225 mm2 silicon wafer, handles token generation with the ultra-low latency Cerebras is known for.

Cerebras Systems logo

Performance And Availability

Together the two compute engines are expected to deliver up to 5x higher tokens per second per kilowatt on Moonshot AI's Kimi 2.6 1T model compared with a Cerebras WSE-only configuration, according to modelling by AMD Performance Labs and Cerebras. Cerebras plans to deploy AMD Helios systems inside its own data centers, and the joint solution will be offered first through Cerebras Cloud in the second half of 2026. "AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," said AMD chair and CEO Dr. Lisa Su.

A Direct Challenge To NVIDIA's Rack

The announcement lands on the same day AMD unveiled the MI450-powered Helios rackscale solution with Anthropic and comes after Cerebras' recent Nasdaq IPO refile. Analysts read the tie-up as an explicit AMD attempt to challenge NVIDIA's Vera Rubin NVL72 rack for the highest-value slice of the inference market, where token latency directly shapes user experience for coding assistants, live agents and voice AI.

Reporting based on coverage from AMD Newsroom, Cerebras press office, HPCwire and Reuters.

Category: Partnerships

Tags: Partnership Semiconductors AI Infrastructure Nvidia AI Chips

Related Articles