Google TPU v6e 32GB vs Google TPU v4 32GB

The TPU v6e delivers 3.3x the BF16 throughput of the TPU v4 (918 vs 275 TFLOPS dense).

TPU v6e: TPU, 2024 TPU v4: TPU, 2020
BF16 dense lead
3.3x
TPU v6e: 918 vs 275 TFLOPS
VRAM
32 vs 32 GB
Same capacity
Bandwidth
1.6 TB/s vs 1.2 TB/s
TPU v6e moves data faster

Specifications

TPU v6e
TPU v4
Architecture
TPUTPU
Launch Year
20242020
Form Factor
VRAM
32 GB32 GB
Memory Bandwidth
1.6 TB/s1.2 TB/s
TDP
Process Node

Performance (TFLOPS)

TPU v6e
TPU v4
BF16
918.0 TFlops 275.0 TFlops
INT8
1836.0 TFlops 275.0 TFlops

FLOPS by Precision

What actually differs

The TPU v6e is the newer part: TPU, launched in 2024, against the TPU v4's TPU from 2020. Newer architectures typically add lower-precision formats and better throughput per watt, so check the precision rows your workload actually uses.

Memory is 32 GB against 32 GB, fed at 1.6 TB/s versus 1.2 TB/s. More memory per GPU means larger models fit before you have to shard across cards, which often matters more than raw TFLOPS for inference.

For LLM training and inference, weight the FP8 and FP16 rows and memory capacity most heavily. For scientific and HPC workloads, the FP64 row is the one to read.

Frequently asked questions

Is the TPU v6e faster than the TPU v4?

At BF16 precision the TPU v6e reaches 918 TFLOPS dense against 275 TFLOPS for the TPU v4. The performance table on this page lists every published precision for both GPUs.

Which has more memory, the TPU v6e or the TPU v4?

Both GPUs carry 32 GB of VRAM. Memory bandwidth is 1.6 TB/s for the TPU v6e and 1.2 TB/s for the TPU v4.

Get Comparison Updates

New GPUs added weekly. Be the first to see how they compare.