Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for Cerebras CS-4 Packs Three Wafer-Scale Chips in a Rack

Cerebras CS-4 Packs Three Wafer-Scale Chips in a Rack

The Cerebras CS-4 puts three WSE-3 Turbo wafers in one rack for 750 PFLOPS, and Cerebras clocks 4,400 tokens per second per user on GPT-OSS-120B.

Dr. Nova Chen
Dr. Nova ChenAug 20, 20265 min read

Cerebras unveiled the CS-4 on August 19, 2026, and it is the first system the company has built around more than one wafer. Three WSE-3 Turbo processors sit in a single rack, wired together on a new modular platform Cerebras calls Nexus, and the whole assembly is aimed squarely at one job: serving large language models to a lot of people at once, very quickly. For a company whose pitch has always been that a dinner-plate-sized chip beats a box full of small ones, this is the moment the argument scales past a single wafer.

  • Each CS-4 rack holds three WSE-3 Turbo wafers for a combined 750 PFLOPS of sparse FP16 compute and 132GB of on-wafer SRAM
  • WSE-3 Turbo keeps the 900,000 cores and 4 trillion transistors of the original WSE-3 but doubles the clock to roughly 2.8 GHz
  • Cerebras reports more than 4,400 tokens per second per user on GPT-OSS-120B, and claims up to 10x more throughput per watt than the CS-3
  • Early access is open to selected customers now, with general availability in Q3 2026

What Is Inside the Cerebras CS-4?

The building block is the WSE-3 Turbo, and the honest description is that it is not a new chip so much as a much faster version of an existing one. It is still built on TSMC 5nm, still carries 900,000 AI-optimized cores and 44GB of SRAM etched directly onto the wafer, still counts about 4 trillion transistors. What changed is the clock, which roughly doubles to 2.8 GHz, and the plumbing that has to feed it. On-wafer memory bandwidth rises to 43.2 PB/s per wafer, and network bandwidth per wafer doubles to 300 GB/s.

Feeding a chip that size is mostly a power and cooling problem, which is where Nexus comes in. Each wafer lives in a self-contained unit Cerebras describes as a backpack, holding its own power delivery, cooling, and I/O. The rack then holds three of those plus power shelves. Splitting compute from power and networking is a maintenance and upgrade story first, but it also means partners can attach the network cards they prefer, and Cerebras has named OpenAI and Amazon Web Services among the partners doing exactly that.

How Fast Is the CS-4 for Inference?

Cerebras puts the rack at 750 PFLOPS of sparse FP16, roughly six times the per-rack performance of the CS-3 generation, and reports throughput above 4,400 tokens per second per user on GPT-OSS-120B. On models beyond 10 trillion parameters, the company cites more than 1,000 tokens per second. Its headline comparison, up to 30x faster than the GPU systems it tested, is a vendor benchmark rather than an independent one, so treat the multiple as directional and wait for third-party numbers.

The per-user framing is the part worth sitting with. Aggregate throughput is what a fleet operator buys; tokens per second per user is what a person waiting on a reasoning model actually feels. Keeping an entire model resident in SRAM instead of streaming weights across HBM is the structural reason wafer-scale designs post such high single-stream numbers, and that advantage grows as models spend more tokens thinking before they answer.

Why Wafer-Scale Matters as Inference Costs Grow

The economics here rhyme with a pattern we have tracked all year. When Etched raised $700M for inference-specific chips, the thesis was that serving, not training, is where compute spending compounds. The software-side version of the same pressure showed up when the Kog inference engine pulled large speedups out of existing GPUs. Custom silicon, better kernels, and wafer-scale integration are three answers to one question: how do you serve far more tokens without buying proportionally more hardware?

Latency between wafers is the interesting engineering detail. In chain topology, Cerebras measures wafer-to-wafer hops at about 2 microseconds, which is what makes three wafers behave like one large accelerator rather than three machines coordinating over a network. Operators who prefer standard networking can run RoCEv2 Ethernet instead.

What to Watch Next

Cerebras has committed publicly to doubling throughput annually through 2029, which is a useful promise to hold the company to. In the meantime, the numbers to watch are independent tokens-per-second-per-dollar figures on current open-weight models, published power draw for a loaded rack, and how quickly general availability turns into deployments outside the partner list.

For readers tracking the compute layer beneath every model launch, follow our artificial intelligence coverage, where serving cost, not training budget, is increasingly the number that decides which AI products ship.

Sources: ServeTheHome — August 19, 2026; The Next Platform — August 19, 2026; Cerebras — August 19, 2026.

More AI Stories