SambaNova logo

SambaNova SambaRack SN40L-16

rack
16× SN40L
Peer-to-peer RDU interconnect with hardware collective offload; PCIe host attach per socket
BF16
10,208
TFLOPS
Power
10.0
kW Total
Memory
1024
GB Total

Sixteen SN40L Reconfigurable Dataflow Units in a single air-cooled rack, which SambaNova positions as able to run models the size of DeepSeek R1 671B and Llama 4 Maverick without liquid cooling. The rack draws an average of 10 kW. Compute is 10.2 PFLOPS of BF16 and memory is three-tier: 8.3 GB of on-chip SRAM, 1 TB of HBM at 28.8 TB/s, and 24 TB of high-capacity DDR. That third tier is the point of the architecture, letting a rack hold many models in DDR and move the active one into HBM in microseconds. Note that the 24 TB of DDR is deliberately not recorded as the system memory figure here, which carries the 1 TB of HBM instead, because the two tiers are different things and conflating them would suggest 24 TB of high-bandwidth memory. SambaNova publishes no rack-level compute figure; the 10.2 PFLOPS is sixteen times the per-socket figure from its own architecture paper, which is safe because its two published rack memory figures of 1 TB and 24 TB are likewise exactly sixteen times the socket. Neither source labels the compute dense or sparse. SambaNova also publishes per-user token rates for this configuration, shown separately below.

We will point you at suppliers who have it. Free, and no signup.

BF16
10,208
TFLOPS

System Details

GPU Configuration

GPU Model: SambaNova SN40L
GPU Count: 16 GPUs
Architecture: RDU
Interconnect: Peer-to-peer RDU interconnect with hardware collective offload; PCIe host attach per socket

System Specifications

Form Factor: rack
Total Power: 10.0 kW
Total Memory: 1.0 TB
On-die SRAM: 8.3 GB
Memory Bandwidth: 28.8 TB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
10,208 TFLOPS 638 TFLOPS —

Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.

Published throughput

Model Tokens/sec Basis Conditions
Llama 3.1 8B 1,042 Per user BF16, 8192 context
Llama 3.1 70B 457 Per user BF16, 8192 context
Llama 3.1 405B 129 Per user BF16, 8192 context

Tokens per second is a benchmark result, not a hardware specification. It depends on the model, the precision and the latency target it was measured under, so figures are only comparable when every condition matches. Vendor claims unless stated otherwise.

Powered by SambaNova SN40L

This system utilizes 16 × SambaNova SN40L GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

—

Per GPU Memory

64 GB

Process Node

5nm

Architecture

RDU

Documentation & Resources

Flopper Spec Sheet

SambaNova SambaRack SN40L-16 specifications, generated from our database. Printable.

Vendor product page

SambaNova SambaRack SN40L-16 documentation

View ↗

Introducing the SN50 RDU, purpose-built for agentic inference (SambaNova blog)

SambaNova • 2026-02-24

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The SambaNova SambaRack SN40L-16 runs 16× SN40L GPUs over Peer-to-peer RDU interconnect with hardware collective offload; PCIe host attach per socket, delivering 10,208 TFLOPS BF16 dense.

© 2026 Flopper.io - Compare the GPUs Powering AI