Cerebras launched the CS-4 on Aug. 18, 2026, a rack-scale AI accelerator that the company says uses 50% fewer components than its previous design. Its Nexus architecture places compute, power conversion, liquid cooling and input-output electronics in self-contained backpacks that plug into the rack. Cerebras plans to use the simpler assembly to speed data-center expansion, but no independent lab has published a hands-on CS-4 benchmark.
What Changed
- CS-4 combines three WSE-3 Turbo wafers in one rack and cuts the component count 50% from Cerebras' prior design.
- Cerebras says modular backpacks reduce installation from days to hours, with first shipments scheduled for the third quarter of 2026.
- Company tests exceeded 4,400 GPT-OSS-120B tokens per second per user; no independent lab has published a hands-on benchmark.
- Hardware revenue fell 23% year over year as cloud and services revenue rose 281%, increasing upfront infrastructure demands.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
Modular rack
Each CS-4 holds three WSE-3 Turbo wafer-scale processors. The Nexus design separates compute, power and input-output hardware into modules that can be manufactured, installed or upgraded without rebuilding the other assemblies. The approach is intended to speed future upgrades, helping power, networking or compute improvements reach customers faster.
A rear-mounted backpack wraps power conversion, direct liquid cooling, control electronics and high-speed connections around each wafer, then attaches vertically to the power array. Cerebras says the change cuts installation from days for its prior system to hours for CS-4 and allows more of the manufacturing work to be automated.
First shipments are scheduled for the third quarter of 2026. Pricing, production volume and delivered customer capacity were not disclosed.
Same silicon, more power
The WSE-3 Turbo is not built on a new process node. It retains the WSE-3's TSMC 5-nanometer process, 46,225-square-millimeter wafer area, four trillion transistors, 900,000 cores and 44 gigabytes of on-chip SRAM.
The performance increase instead comes from driving the existing silicon harder. Cerebras moved power conversion closer to the processor, reducing board-level losses and allowing twice as much power to reach the wafer. That permits a higher operating frequency, although the company did not disclose the clock speed.
The new input-output module supports standard Ethernet connections as well as switch-free Direct Wafer Links. Cerebras says those direct links reduce wafer-to-wafer latency from five microseconds on WSE-3 to as little as two microseconds on WSE-3 Turbo.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
Performance and power limits
In testing disclosed Aug. 18, 2026, Cerebras measured more than 4,400 GPT-OSS-120B output tokens per second per user and said CS-4 ran as much as 30 times faster than GPU services. Cerebras produced the CS-4 result, while Artificial Analysis supplied the external GPU baseline. The comparison can change with model architecture, context length, precision and serving configuration.
The headline compute figure also depends on sparsity, a method that skips some calculations. An independent technical analysis estimated dense FP16 performance at 25 petaflops per wafer, one-tenth of the 250 sparse-petaflops figure promoted for WSE-3 Turbo, and noted that sparsity generally does not help large-language-model inference.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Cerebras did not disclose full-rack power draw. The same analysis estimated 120 to 140 kilowatts for a fully populated three-wafer rack.
Deployment bill
Cerebras' hardware revenue fell 23% from a year earlier to $54.12 million in the second quarter ended June 30, 2026, from $70.3 million. Cloud and services revenue rose 281% to $125.99 million from $33.03 million. That shift requires Cerebras to fund, install and operate more data-center equipment before collecting service revenue over time.
The company targets 600 megawatts of deployed computing capacity by the end of 2027. It also says its hardware will run four times faster and deliver 20 times more throughput by then, compared with August 2026 levels. Those are forward-looking targets, not measured CS-4 results.
Paul Meeks, head of technology research at Freedom Capital Markets, said, "There might be a little bit more of a gap between the buildout and the revenue than people expect."
Frequently Asked Questions
What is the Cerebras CS-4?
CS-4 is a rack-scale AI accelerator that holds three WSE-3 Turbo wafer-scale processors. Its Nexus design separates compute, power and input-output hardware into modules intended to speed deployment and upgrades.
How is WSE-3 Turbo different from WSE-3?
WSE-3 Turbo keeps the same TSMC 5-nanometer process, four trillion transistors, 900,000 cores and 44 gigabytes of SRAM. Cerebras delivers more power to the wafer to run it at a higher frequency, though it did not disclose the clock speed.
How fast is CS-4?
Cerebras measured more than 4,400 GPT-OSS-120B output tokens per second per user and said CS-4 was up to 30 times faster than GPU services. The CS-4 result was company-produced, and no independent lab has published a hands-on benchmark.
How much power does a CS-4 rack use?
Cerebras did not disclose full-rack power draw. An independent technical analysis estimated 120 to 140 kilowatts for a fully populated three-wafer rack.
When will CS-4 ship, and what will it cost?
First shipments are scheduled for the third quarter of 2026. Cerebras did not disclose pricing, production volume or delivered customer capacity.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR