Upgrade to SI Premium - Free Trial

Cerebras launches CS-4 AI accelerator built on three wafer-scale chips

August 19, 2026 6:00 AM

Cerebras Systems (NASDAQ: CBRS) introduced the CS-4, a rack-scale AI accelerator built from three of its newly announced Wafer Scale Engine 3 Turbo (WSE-3T) processors, according to a company statement. First shipments are scheduled to begin this quarter.

The CS-4 delivers 750 petaflops of AI compute, 129.6 petabytes per second of memory bandwidth, and 7.2 terabits per second of system I/O bandwidth. The company claims the system achieves more than 4,400 tokens per second per user on the GPT-OSS-120B model, which it says is up to 30 times faster than GPU-based solutions and twice as fast as its previous CS-3 system.

The CS-4 also claims up to 10x more throughput per watt compared to the CS-3. The I/O latency is reduced from five microseconds on the CS-3 to as low as two microseconds, enabling clusters capable of supporting models with more than 50 trillion parameters.

The underlying WSE-3T processor contains four trillion transistors and 900,000 cores across 46,225 square millimeters of silicon, with 44GB of on-wafer SRAM. Each WSE-3T delivers 250 petaflops of AI compute and 43.2 petabytes per second of memory bandwidth, doubling the figures of the prior WSE-3.

The CS-4 introduces a new modular system design called the Nexus Platform Architecture. A rear-mounted "backpack" integrates power conversion, liquid cooling, high-speed I/O, and control electronics around each wafer. Cerebras says the design reduces deployment time from days to hours and uses 50% fewer components with 60% more automated manufacturing compared to the prior generation.

The system supports RoCE v2 RDMA over Ethernet and a new Direct Wafer Links mode for switch-free wafer-to-wafer communication. Cerebras said the I/O subsystem is compatible with disaggregated inference configurations involving partners including AMD Helios and AWS Trainium.

Categories

Corporate News Hot Corp. News

Next Articles