| Precision | Peak (PFLOPS) | Per GPU (TFLOPS) | TFLOPS/W |
|---|---|---|---|
| BF16 | 4112.6 | 459.0 | — |
| FP8 | 4112.6 | 459.0 | — |
All figures are dense. Vendors commonly headline the sparse number, which is twice the dense one.
3D torus ICI, 1,200 GBps bidirectional per chip, 50 Gbps data center network per chip
A full TPU v5p pod is 8,960 chips in a 3D torus, the largest TPU pod Google built before Ironwood. Per chip Google publishes 459 TFLOPS bf16 and 459 TFLOPS fp8 as two separate rows carrying the same number: v5p gets no fp8 throughput advantage over bf16. Across the pod that is 4,112.6 PFLOPS at either precision. Memory is 8,960 x 95 GiB HBM at 2,765 GBps per chip, with 1,200 GBps of bidirectional inter-chip interconnect and 50 Gbps of data center network per chip, and two TensorCores per chip. The table also lists four SparseCores per chip; those are embedding-lookup dataflow processors and have nothing to do with weight sparsity, so no figure here is a sparse figure. Google publishes no pod-level compute or power figure; the compute above is chip count times published per-chip peak.