Google TPU v5p Pod
A full TPU v5p pod is 8,960 chips in a 3D torus, the largest TPU pod Google built before Ironwood. Per chip Google publishes 459 TFLOPS bf16 and 459 TFLOPS fp8 as two separate rows carrying the same number: v5p gets no fp8 throughput advantage over bf16. Across the pod that is 4,112.6 PFLOPS at either precision. Memory is 8,960 x 95 GiB HBM at 2,765 GBps per chip, with 1,200 GBps of bidirectional inter-chip interconnect and 50 Gbps of data center network per chip, and two TensorCores per chip. The table also lists four SparseCores per chip; those are embedding-lookup dataflow processors and have nothing to do with weight sparsity, so no figure here is a sparse figure. Google publishes no pod-level compute or power figure; the compute above is chip count times published per-chip peak.
We will point you at suppliers who have it. Free, and no signup.
System Details
GPU Configuration
System Specifications
Precision Performance Breakdown
| Precision | System Performance | Per GPU | Efficiency |
|---|---|---|---|
| 4112.640 PFLOPS | 459.0 TFLOPS | — | |
| 4112.640 PFLOPS | 459.0 TFLOPS | — |
All figures are dense. Vendors commonly headline the number, which is twice the dense one.
Powered by Google TPU v5p
This system utilizes 8960 × Google TPU v5p 95GB GPUs, each delivering exceptional performance for AI and HPC workloads.
Per GPU TDP
450W
Per GPU Memory
95 GB
Process Node
5nm
Architecture
TPU v5p
Documentation & Resources
Official Datasheet
Google TPU v5p Pod technical specifications
Cloud TPU v5p
Google • 2023-12-07
Typical Use Cases
The Google TPU v5p Pod runs 8960× TPU v5p GPUs over 3D torus ICI, 1,200 GBps bidirectional per chip, 50 Gbps data center network per chip, delivering 4,113 PFLOPS FP8 dense.