Existing Confirmed

Paper on AFM-server

Google • United States of America • June 10, 2024
Total GPUs
8.2K
H100 Equiv
1.1K
Power
2.8 MW
FP16
2.3K PF

Specifications

Primary GPU Google TPU v4
GPU Count 8.2K
Sector Private

Additional Information

"We train AFM-server from scratch for 6.3T tokens on 8192 TPUv4 chips, using a sequence length of 4096 and a batch-size of 4096 sequences." "The AFM models are pre-trained on v4 and v5p Cloud TPU clusters" "AFM-server was trained on 8192 TPUv4 chips provisioned as 8 × 1024 chip slices, where slices are connected together by the data-center network (DCN)"

Cloud TPU clusters implies they trained on Google infrastructure