NVIDIA GB200 Grace Blackwell Superchip 372GB

Blackwell Superchip 2024 4nm Spec confidence: Official
FP8 (dense)
10,000
TFLOPS
FP32
160
TFLOPS
Memory
372 GB
HBM3e
Bandwidth
16.0 TB/s
memory

Overview

The NVIDIA GB200 Grace Blackwell Superchip is the building block of the GB200 NVL72 rack: one 72-core Grace CPU joined to two B200 GPUs over NVLink-C2C at 900 GB/s, with 372 GB of HBM3e at 16 TB/s across the pair and 480 GB of LPDDR5X on the CPU side. Every performance figure on this page is for the whole superchip, so it is twice a single Blackwell GPU: 20 PFLOPS of dense NVFP4, 10 PFLOPS of dense FP8, 160 TFLOPS of FP32. Two things follow from the packaging rather than the silicon. The GPUs in a superchip are binned slightly above a standalone B200, so per-GPU dense NVFP4 is 10 PFLOPS here against 9 on an HGX B200 board. And because the CPU is coherently attached rather than sitting behind PCIe, a model can address CPU memory at NVLink speed instead of being staged across a bus, which is the practical reason 72-GPU racks are built from these rather than from discrete GPUs and separate hosts. Thirty-six superchips make an NVL72.

Performance

Peak theoretical throughput by precision type

PrecisionDense2:4 Sparse
FP64
64-bit floating point
80TFLOPS Structured sparsity is a tensor-core feature; this vector precision has no sparse form
FP32
32-bit floating point
160TFLOPS Structured sparsity is a tensor-core feature; this vector precision has no sparse form
TF32
TensorFloat-32
2,500TFLOPS 5,000TFLOPS
BF16
Brain Float 16
5,000TFLOPS 10,000TFLOPS
FP16
16-bit floating point
5,000TFLOPS 10,000TFLOPS
FP8
8-bit floating point
10,000TFLOPS 20,000TFLOPS
FP6
10,000TFLOPS 20,000TFLOPS
FP4
4-bit floating point
20,000TFLOPS 40,000TFLOPS
INT8
8-bit integer
10,000TOPS 20,000TOPS

Specifications

Architecture

Blackwell

Form Factor

Superchip

Launch Year

2024

Process Node

4nm

Memory

372 GB HBM3e

Bandwidth

16,000 GB/s

TDP

No verified data available

Max power

No verified data available

Interconnect

900 GB/s NVLink-C2C

bidirectional, per GPU

Spec Confidence

Official

Full Specifications

Compute Engine
GPUs per Module 2 (all figures are totals for the module)
Memory
Memory 372 GB
Memory Type HBM3e
Bandwidth 16.0 TB/s
Interface Width No verified data available
Interconnect & I/O
GPU-to-GPU NVLink-C2C
Interconnect Bandwidth 900 GB/s bidirectional, per GPU
Power & Thermal
TDP No verified data available
Max power No verified data available
Enterprise Features
Sparsity Yes
General
Form Factor Superchip
Architecture Blackwell
Process Node 4nm
Launch Year 2024

Datasheet & Resources

Data Provenance

Every figure traced to a source

Primary Source

Publisher
NVIDIA
Published
2024-03-18

Data Quality

Spec confidence
Official
Clock basis
Boost
Core precisions with figures
9 of 9
Normalization
All values in TFLOPS

Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.

Browse GPU Cloud Providers

Similar GPUs

Frequently Asked Questions

How many TFLOPS does the NVIDIA GB200 have?

The NVIDIA GB200 delivers 160 TFLOPS FP32, 5,000 TFLOPS FP16 and 10,000 TFLOPS FP8 at peak.

What is the power consumption of the NVIDIA GB200?

Flopper does not currently have a verified TDP (Thermal Design Power) figure for the NVIDIA GB200.

How much memory does the NVIDIA GB200 have?

The NVIDIA GB200 is equipped with 372 GB of memory with 16,000 GB/s of memory bandwidth.

What architecture is the NVIDIA GB200 based on?

The NVIDIA GB200 is based on the Blackwell architecture, launched in 2024.

Stay Updated on GPU Releases

Get notified when new GPUs are added or specifications are updated.

Loading verification...

No spam, unsubscribe anytime.

Back to GPUs