Cerebras Systems unveiled the CS-4 rack-scale AI accelerator Tuesday, claiming up to 30 times faster inference than GPU-based systems and doubling the performance of its predecessor.
Cerebras Systems unveiled the CS-4 rack-scale AI accelerator Tuesday, claiming up to 30 times faster inference than GPU-based systems and doubling the performance of its predecessor.

Cerebras Systems launched the CS-4 AI accelerator Tuesday, claiming up to 30 times faster inference than GPU solutions with 750 PFLOPs of compute, a direct challenge to Nvidia's grip on the AI inference market.
"Being 30 times faster doesn't just make a response feel fast. It gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use in the same wall-clock time," Sean Lie, CTO and co-founder of Cerebras, said.
The rack-scale CS-4 pairs three WSE-3 Turbo wafers — each containing four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon — with 129.6 petabytes per second of memory bandwidth and 7.2 terabits per second of I/O. On GPT-OSS-120B, the system delivers more than 4,400 tokens per second per user, up to 30 times faster than GPU solutions, and supports models exceeding 50 trillion parameters.
The launch lands one week after Cerebras reported Q2 core revenue of $209.9 million, up 103% year over year, and $25.4 billion in remaining performance obligations. CBRS shares, which have swung between $169 and $265 over the past month, closed near $219 after the earnings selloff, with UBS, Morgan Stanley, and Needham maintaining bullish ratings.
WSE-3 Turbo doubles compute per wafer
The WSE-3 Turbo doubles AI compute to 250 PFLOPs per wafer and memory bandwidth to 43.2 petabytes per second compared with the WSE-3. On-chip fabric bandwidth rises to 53.5 petabytes per second, while I/O latency drops from five microseconds to as low as two microseconds. The CS-4 also delivers up to 10 times more throughput per watt than the CS-3, a metric that directly affects data center operating costs.
The new Nexus platform architecture separates compute, power, and I/O into modular elements. Power conversion moves from roughly 50 millimeters away from the processor on conventional GPU boards to approximately 0.5 millimeters, cutting board-level power loss and enabling higher operating frequencies. A pluggable "backpack" design folds power conversion, liquid cooling, and control electronics into a self-contained unit, reducing deployment time from days to hours.
AMD Helios pairing targets 5x efficiency gain
The CS-4's programmable I/O subsystem supports standards-based RoCE v2 RDMA over Ethernet, allowing integration with heterogeneous systems. This enables disaggregated inference architectures in which a purpose-built prefill engine processes prompts before handing off to Cerebras for ultra-low-latency decode. Cerebras has already partnered with AMD on such a configuration, pairing AMD's Helios Rackscale solutions with the WSE for up to five times more tokens per second per watt than a Cerebras-only setup. AWS Trainium is also named as a partner.
Dylan Patel, founder and CEO of SemiAnalysis, said the CS-4's improvements in deployability and networking enable scaling to larger models for "large scale token factories."
Investor implications
Cerebras' competitive positioning hinges on whether its performance claims hold up in independent benchmarks. The company did not disclose the specific GPU configuration used for the 30x comparison, noting only that "actual throughput varies by model architecture, context length, precision, and serving configuration." If confirmed, the efficiency advantage could shift procurement decisions at hyperscale data center operators that currently deploy Nvidia's H100 and B200 accelerators.
The company's $25.4 billion in remaining performance obligations, anchored by a $10 billion-plus OpenAI contract for 750 megawatts of inference capacity, provides contracted revenue visibility through 2028. First CS-4 shipments begin this quarter, with the company guiding to full-year 2026 core revenue of $880–890 million and a tripling of core revenue in 2027.
CBRS trades at roughly 50 times trailing sales, compressing to about 16 times if the 2027 guidance is delivered. The November 9 lock-up expiration for CEO Feldman and CTO Lie's shares represents a supply overhang that investors will need to monitor.
This article is for informational purposes only and does not constitute investment advice.