Google TPU v4 Pod

pod
4096× TPU v4
3D mesh ICI, 1.1 PB/s all-reduce bandwidth per pod, 24 TB/s bisection bandwidth
2020
BF16
1126.4
PFLOPS
INT8
1126.4
PFLOPS
Power
kW Total
Memory
134218
GB Total

A full TPU v4 pod is 4,096 chips in a 3D mesh. Google publishes a single peak compute figure of 275 teraflops per chip covering both bf16 and int8 rather than separate rows, giving 1,126.4 PFLOPS across the pod; Google prints this rounded as "1.1 exaflops (bf16 or int8)". Also published per pod: 1.1 PB/s all-reduce bandwidth and 24 TB/s bisection bandwidth. Memory is 4,096 x 32 GiB HBM2 at 1,200 GBps per chip. Google measures per-chip power at 90 W minimum, 170 W mean and 192 W maximum, but publishes no pod-level power, so total system power is left blank here rather than extrapolated from silicon alone. Google publishes no sparse figures for any TPU, so all figures are dense.

We will point you at suppliers who have it. Free, and no signup.

BF16
1126.40
PFLOPS
INT8
1126.40
PFLOPS

System Details

GPU Configuration

GPU Count: 4096 GPUs
Architecture: TPU
Interconnect: 3D mesh ICI, 1.1 PB/s all-reduce bandwidth per pod, 24 TB/s bisection bandwidth

System Specifications

Form Factor: pod
Total Power:
Total Memory: 134218 GB
Memory Bandwidth: 4915200 GB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
1126.400 PFLOPS 275.0 TFLOPS
1126.400 PFLOPS 275.0 TFLOPS

All figures are dense. Vendors commonly headline the number, which is twice the dense one.

Powered by Google TPU v4

This system utilizes 4096 × Google TPU v4 32GB GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

Per GPU Memory

32 GB

Process Node

Architecture

TPU

Documentation & Resources

Official Datasheet

Google TPU v4 Pod technical specifications

Download PDF

Cloud TPU v5p

Google • 2023-12-07

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The Google TPU v4 Pod runs 4096× TPU v4 GPUs over 3D mesh ICI, 1.1 PB/s all-reduce bandwidth per pod, 24 TB/s bisection bandwidth, delivering 1,126 PFLOPS BF16 dense.