Cerebras Systems on Tuesday unveiled the CS-4, its first multi-wafer rack-scale AI system and the debut of a new architecture called Nexus that it is pitching as an inference machine for frontier models.
Three WSE-3 Turbo Wafers In One Rack
The CS-4 packs three WSE-3 Turbo processors, each fabricated by TSMC on a 5nm node with 4 trillion transistors, 900,000 AI-optimized cores, and 44 GB of SRAM on the wafer itself. Cerebras says the three-wafer configuration delivers 750 PFLOPS of AI compute, 129.6 PB/s of memory bandwidth, and 7.2 Tb/s of I/O — enough, the company claims, to run models with more than 50 trillion parameters when systems are chained together over Direct Wafer Links.
Speed Over New Silicon
Notably, the underlying chip is not new. Cerebras extracted the CS-4's headline performance by pairing the existing WSE-3 in a Turbo variant with the new Nexus rack architecture, modular Wafer-Scale Backpacks that integrate power conversion, liquid cooling and I/O, and a switchless direct-wafer interconnect that drops inter-wafer latency from about five microseconds to two. In a public test on OpenAI's GPT-OSS-120B, Cerebras said the CS-4 processed more than 4,400 tokens per second per user — a figure it claims is up to 30x faster than GPU-based deployments in certain configurations.
The Inference Land-Grab
The launch lands as AI economics tilt from training toward inference, where coding agents, reasoning systems and voice apps are pushing token throughput to the top of every buyer's spec sheet. Cerebras is already benefiting: last week the company reported Q2 cloud revenue up 281% as its OpenAI GPT-5.6 Sol deployment ramped, and its AMD Helios inference partnership aims to bring wafer-scale decode to a broader stack. CEO Andrew Feldman told analysts Cerebras expects to deliver 600 MW of compute capacity by year-end 2027, with next-generation systems planned to be four times faster and 20x higher throughput. The first CS-4 units ship in the third quarter; Cerebras did not disclose pricing.
Reporting based on coverage from Techzine, Reuters and Cerebras's investor release.
