Cerebras launched the CS-4 on Aug. 18, 2026, a rack-scale AI accelerator that the company says uses 50% fewer components than its previous design. Its Nexus architecture places compute, power conversion, liquid cooling and input-output electronics in self-contained backpacks that plug into the rack. Cerebras plans to use the simpler assembly to speed data-center expansion, but no independent lab has published a hands-on CS-4 benchmark.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Modular rack

Each CS-4 holds three WSE-3 Turbo wafer-scale processors. The Nexus design separates compute, power and input-output hardware into modules that can be manufactured, installed or upgraded without rebuilding the other assemblies. The approach is intended to speed future upgrades, helping power, networking or compute improvements reach customers faster.

A rear-mounted backpack wraps power conversion, direct liquid cooling, control electronics and high-speed connections around each wafer, then attaches vertically to the power array. Cerebras says the change cuts installation from days for its prior system to hours for CS-4 and allows more of the manufacturing work to be automated.

First shipments are scheduled for the third quarter of 2026. Pricing, production volume and delivered customer capacity were not disclosed.

Same silicon, more power

The WSE-3 Turbo is not built on a new process node. It retains the WSE-3's TSMC 5-nanometer process, 46,225-square-millimeter wafer area, four trillion transistors, 900,000 cores and 44 gigabytes of on-chip SRAM.

The performance increase instead comes from driving the existing silicon harder. Cerebras moved power conversion closer to the processor, reducing board-level losses and allowing twice as much power to reach the wafer. That permits a higher operating frequency, although the company did not disclose the clock speed.

The new input-output module supports standard Ethernet connections as well as switch-free Direct Wafer Links. Cerebras says those direct links reduce wafer-to-wafer latency from five microseconds on WSE-3 to as little as two microseconds on WSE-3 Turbo.

Performance and power limits

In testing disclosed Aug. 18, 2026, Cerebras measured more than 4,400 GPT-OSS-120B output tokens per second per user and said CS-4 ran as much as 30 times faster than GPU services. Cerebras produced the CS-4 result, while Artificial Analysis supplied the external GPU baseline. The comparison can change with model architecture, context length, precision and serving configuration.

The headline compute figure also depends on sparsity, a method that skips some calculations. An independent technical analysis estimated dense FP16 performance at 25 petaflops per wafer, one-tenth of the 250 sparse-petaflops figure promoted for WSE-3 Turbo, and noted that sparsity generally does not help large-language-model inference.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Cerebras did not disclose full-rack power draw. The same analysis estimated 120 to 140 kilowatts for a fully populated three-wafer rack.

Deployment bill

Cerebras' hardware revenue fell 23% from a year earlier to $54.12 million in the second quarter ended June 30, 2026, from $70.3 million. Cloud and services revenue rose 281% to $125.99 million from $33.03 million. That shift requires Cerebras to fund, install and operate more data-center equipment before collecting service revenue over time.

The company targets 600 megawatts of deployed computing capacity by the end of 2027. It also says its hardware will run four times faster and deliver 20 times more throughput by then, compared with August 2026 levels. Those are forward-looking targets, not measured CS-4 results.

Paul Meeks, head of technology research at Freedom Capital Markets, said, "There might be a little bit more of a gap between the buildout and the revenue than people expect."

Frequently Asked Questions

What is the Cerebras CS-4?

CS-4 is a rack-scale AI accelerator that holds three WSE-3 Turbo wafer-scale processors. Its Nexus design separates compute, power and input-output hardware into modules intended to speed deployment and upgrades.

How is WSE-3 Turbo different from WSE-3?

WSE-3 Turbo keeps the same TSMC 5-nanometer process, four trillion transistors, 900,000 cores and 44 gigabytes of SRAM. Cerebras delivers more power to the wafer to run it at a higher frequency, though it did not disclose the clock speed.

How fast is CS-4?

Cerebras measured more than 4,400 GPT-OSS-120B output tokens per second per user and said CS-4 was up to 30 times faster than GPU services. The CS-4 result was company-produced, and no independent lab has published a hands-on benchmark.

How much power does a CS-4 rack use?

Cerebras did not disclose full-rack power draw. An independent technical analysis estimated 120 to 140 kilowatts for a fully populated three-wafer rack.

When will CS-4 ship, and what will it cost?

First shipments are scheduled for the third quarter of 2026. Cerebras did not disclose pricing, production volume or delivered customer capacity.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Anthropic Weighs Microsoft Maia Chip Deal as Claude Demand Grows
Anthropic is in early talks to rent servers powered by Microsoft’s Maia AI chips, The Information reported Thursday, citing two people familiar with the discussions. The talks would give the Claude ma
Nvidia Didn't Just Launch Chips at GTC. It Launched a Lock-In Machine.
Monday at the SAP Center in San Jose, Jensen Huang held up a chip. Rotated it under the stage lights, slow, deliberate, the way he always does. A jeweler showing off a diamond. Thirty thousand people
Nvidia still wins per chip. Google just changed what counts.
Google Cloud on Wednesday unveiled two new TPUs at Cloud Next 2026, splitting its eighth-generation design into a training chip and an inference chip for the first time in the program's decade-long hi
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai