CS

Cerebras CS-3 Inference Cluster (88-node)

cluster
88× WSE-3
On-wafer fabric per engine, 214 Pb/s; cluster networking via dedicated NET cabinets
2024
Power
2540.0
kW Total
Memory
GB Total

The largest Cerebras inference cluster configuration published, at eighty-eight nodes provisioned as fifty-two cabinets: forty-four WSE cabinets holding two CS-3 systems each, plus eight NET cabinets. Provisioned heat load is approximately 2.54 MW, about 2.4 MW on water and 0.14 MW on air, Cerebras's own provisioned total including networking. Fully loaded a WSE cabinet weighs 843 kg and a NET cabinet 963 kg. The cluster holds 3.872 TB of on-chip SRAM and no DRAM at all. Each engine is rated at 125 petaFLOPS, footnoted by Cerebras as sparse; because that sparsity is unstructured rather than 2:4 structured, no dense equivalent can be derived and none is published.

We will point you at suppliers who have it. Free, and no signup.

System Details

GPU Configuration

GPU Model: Cerebras WSE-3
GPU Count: 88 GPUs
Architecture: Wafer Scale Engine 3
Interconnect: On-wafer fabric per engine, 214 Pb/s; cluster networking via dedicated NET cabinets

System Specifications

Form Factor: cluster
Total Power: 2540.0 kW
Total Memory:
Memory Bandwidth: 1848000000 GB/s

Powered by Cerebras WSE-3

This system utilizes 88 × Cerebras WSE-3 GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

Per GPU Memory

Process Node

TSMC 5nm

Architecture

Wafer Scale Engine 3

Documentation & Resources

Official Datasheet

Cerebras CS-3 Inference Cluster (88-node) technical specifications

Download PDF

Cerebras CS-3 datasheet: system, rack and inference cluster specifications

Cerebras • Latest version

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The Cerebras CS-3 Inference Cluster (88-node) runs 88× WSE-3 GPUs over On-wafer fabric per engine, 214 Pb/s; cluster networking via dedicated NET cabinets.