| Precision | Peak (TFLOPS) | Per GPU (TFLOPS) | TFLOPS/W |
|---|---|---|---|
| FP64 | 10 | 1.3 | — |
| FP32 | 600 | 75 | — |
| TF32 | 9,000 | 1,125 | — |
| FP16 | 18,000 | 2,250 | — |
| BF16 | 18,000 | 2,250 | — |
| FP8 | 36,000 | 4,500 | — |
| FP4 | 108,000 | 13,500 | — |
| FP6 | 36,000 | 4,500 | — |
All figures are dense. Vendors commonly headline the sparse number, which is twice the dense one.
NVSwitch + NVLink
Eight Blackwell Ultra SXM GPUs on the HGX B300 baseboard with Intel Xeon 6776P processors, in a 10U chassis. NVIDIA quotes 144 PFLOPS of FP4 inference performance for this machine, which is the sparse figure; the dense equivalent is 108 PFLOPS, and both are recorded. Power is given by NVIDIA as approximately 14 kW.