AWS

AWS EC2 trn2.48xlarge

rack
16× Trainium2
NeuronLink-v3 at 1,024 GB/s per chip within the instance; EFAv3 networking at 3,200 Gbps
2024
FP32
2.9
PFLOPS
TF32
10.7
PFLOPS
Power
kW Total
Memory
1573
GB Total

The trn2.48xlarge is AWS's single-node Trainium2 instance: sixteen Trainium2 chips joined by NeuronLink-v3 at 1,024 GB/s per chip, with 1,536 GiB of device memory at 46.4 TB/s, 192 vCPUs, 2 TB of host memory and 3,200 Gbps of EFAv3 networking. AWS rates it at 20.8 PFLOPS of dense FP8, 10.7 PFLOPS across dense FP16, BF16 and TF32, and 2.9 PFLOPS of dense FP32, with a separate sparse figure of 41 PFLOPS covering FP8, FP16, BF16 and TF32 but not FP32. Note that the sparse figure is larger than the dense FP8 one, so the two must not be confused: dense FP16 is 10.7 PFLOPS, not 41. AWS's sparse mode runs at the chip's low-precision rate whatever format is fed in, which is why 41 is 1.97x the dense FP8 rate but 3.84x the dense BF16 rate; it is not 2:4 structured sparsity. AWS publishes no power figure for this instance or for the Trainium2 chip, so total system power is left blank rather than estimated. The trn2u.48xlarge variant carries the same specification and is the instance from which UltraServers are composed.

We will point you at suppliers who have it. Free, and no signup.

FP32
2.90
PFLOPS
TF32
10.70
PFLOPS
FP16
10.70
PFLOPS
BF16
10.70
PFLOPS

System Details

GPU Configuration

GPU Model: AWS Trainium2
GPU Count: 16 GPUs
Architecture: NeuronCore-v3
Interconnect: NeuronLink-v3 at 1,024 GB/s per chip within the instance; EFAv3 networking at 3,200 Gbps

System Specifications

Form Factor: rack
Total Power:
Total Memory: 1573 GB
Memory Bandwidth: 46400 GB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
2.900 PFLOPS 181.3 TFLOPS
10.700 PFLOPS 668.8 TFLOPS
10.700 PFLOPS 668.8 TFLOPS
10.700 PFLOPS 668.8 TFLOPS
20.800 PFLOPS 1300.0 TFLOPS

All figures are dense. Vendors commonly headline the number, which is twice the dense one.

Powered by AWS Trainium2

This system utilizes 16 × AWS Trainium2 GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

Per GPU Memory

96 GB

Process Node

Architecture

NeuronCore-v3

Documentation & Resources

Official Datasheet

AWS EC2 trn2.48xlarge technical specifications

Download PDF

Trainium2 Architecture (AWS Neuron Documentation)

AWS • Latest version

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The AWS EC2 trn2.48xlarge runs 16× Trainium2 GPUs over NeuronLink-v3 at 1,024 GB/s per chip within the instance; EFAv3 networking at 3,200 Gbps, delivering 20.8 PFLOPS FP8 dense.