SA

SambaNova SambaRack SN40L-16

rack
16× SN40L
Peer-to-peer RDU interconnect with hardware collective offload; PCIe host attach per socket
BF16
10.2
PFLOPS
Power
10.0
kW Total
Memory
1049
GB Total

Sixteen SN40L Reconfigurable Dataflow Units in a single air-cooled rack, which SambaNova positions as able to run models the size of DeepSeek R1 671B and Llama 4 Maverick without liquid cooling. The rack draws an average of 10 kW. Compute is 10.2 PFLOPS of BF16 and memory is three-tier: 8.3 GB of on-chip SRAM, 1 TB of HBM at 28.8 TB/s, and 24 TB of high-capacity DDR. That third tier is the point of the architecture, letting a rack hold many models in DDR and move the active one into HBM in microseconds. Note that the 24 TB of DDR is deliberately not recorded as the system memory figure here, which carries the 1 TB of HBM instead, because the two tiers are different things and conflating them would suggest 24 TB of high-bandwidth memory. SambaNova publishes no rack-level compute figure; the 10.2 PFLOPS is sixteen times the per-socket figure from its own architecture paper, which is safe because its two published rack memory figures of 1 TB and 24 TB are likewise exactly sixteen times the socket. Neither source labels the compute dense or sparse. SambaNova also publishes per-user token rates for this configuration, shown separately below.

We will point you at suppliers who have it. Free, and no signup.

BF16
10.21
PFLOPS

System Details

GPU Configuration

GPU Model: SambaNova SN40L
GPU Count: 16 GPUs
Architecture: RDU
Interconnect: Peer-to-peer RDU interconnect with hardware collective offload; PCIe host attach per socket

System Specifications

Form Factor: rack
Total Power: 10.0 kW
Total Memory: 1049 GB
Memory Bandwidth: 28800 GB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
10.208 PFLOPS 638.0 TFLOPS

All figures are dense. Vendors commonly headline the number, which is twice the dense one.

Powered by SambaNova SN40L

This system utilizes 16 × SambaNova SN40L GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

Per GPU Memory

64 GB

Process Node

5nm

Architecture

RDU

Documentation & Resources

Official Datasheet

SambaNova SambaRack SN40L-16 technical specifications

Download PDF

SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts (Prabhakar et al., SambaNova Systems) -- Table II carries the chip parameters SambaNova publishes nowhere else

SambaNova • Latest version

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The SambaNova SambaRack SN40L-16 runs 16× SN40L GPUs over Peer-to-peer RDU interconnect with hardware collective offload; PCIe host attach per socket, delivering 10.2 PFLOPS BF16 dense.