AWS logo

AWS Trn2 UltraServer

rack
64× Trainium2
NeuronLink-v3 at 1,024 GB/s per chip within each instance and 256 GB/s per chip between instances; EFAv3 networking at 3,200 Gbps
2024
FP32
11.6
PFLOPS
TF32
42.8
PFLOPS
Power
—
kW Total
Memory
6144
GB Total

The Trn2 UltraServer joins four trn2u.48xlarge instances into a single 64-chip Trainium2 machine, connected by NeuronLink-v3 at 1,024 GB/s per chip within an instance and 256 GB/s per chip between them. It carries 6,144 GiB of device memory at 185.6 TB/s, 768 vCPUs and 8 TB of host memory. AWS rates it at 83.2 PFLOPS of dense FP8, 42.8 PFLOPS across dense FP16, BF16 and TF32, and 11.6 PFLOPS of dense FP32, with a separate sparse figure of 164 PFLOPS covering FP8, FP16, BF16 and TF32 but not FP32. The sparse figure exceeds the dense FP8 one, so the two must not be conflated: dense FP16 is 42.8 PFLOPS, not 164. AWS's sparse mode runs at the chip's low-precision rate regardless of input format, making the sparse figure 1.97x dense FP8 but 3.84x dense BF16, so it is not 2:4 structured sparsity. Every figure here is AWS's own published aggregate and each one reconciles exactly against the Trainium2 chip specification. AWS publishes no power figure for the UltraServer or the chip.

We will point you at suppliers who have it. Free, and no signup.

FP32
11.6
PFLOPS
TF32
42.8
PFLOPS
FP16
42.8
PFLOPS
BF16
42.8
PFLOPS

System Details

GPU Configuration

GPU Model: AWS Trainium2
GPU Count: 64 GPUs
Architecture: NeuronCore-v3
Interconnect: NeuronLink-v3 at 1,024 GB/s per chip within each instance and 256 GB/s per chip between instances; EFAv3 networking at 3,200 Gbps

System Specifications

Form Factor: rack
Total Power: —
Total Memory: 6.1 TB
Memory Bandwidth: 185.6 TB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
11.6 PFLOPS 181 TFLOPS —
42.8 PFLOPS 669 TFLOPS —
42.8 PFLOPS 669 TFLOPS —
42.8 PFLOPS 669 TFLOPS —
83.2 PFLOPS 1,300 TFLOPS —

Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.

Powered by AWS Trainium2

This system utilizes 64 × AWS Trainium2 GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

—

Per GPU Memory

96 GB

Process Node

—

Architecture

NeuronCore-v3

Documentation & Resources

Flopper Spec Sheet

AWS Trn2 UltraServer specifications, generated from our database. Printable.

Vendor product page

AWS Trn2 UltraServer documentation

View ↗

AWS announces Amazon EC2 Inf2 instances (Preview)

AWS • 2022-11-29

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The AWS Trn2 UltraServer runs 64× Trainium2 GPUs over NeuronLink-v3 at 1,024 GB/s per chip within each instance and 256 GB/s per chip between instances; EFAv3 networking at 3,200 Gbps, delivering 83.2 PFLOPS FP8 dense.

© 2026 Flopper.io - Compare the GPUs Powering AI