AWS

AWS Trn3 UltraServer

rack
144× Trainium3
NeuronSwitch-v1 all-to-all fabric over NeuronLink-v4, 2 TB/s per chip
2025
FP32
26.4
PFLOPS
TF32
96.6
PFLOPS
Power
kW Total
Memory
21234
GB Total

The Trn3 UltraServer is AWS's rack-scale Trainium3 machine, generally available since 2 December 2025 and reaching 144 chips joined by NeuronSwitch-v1, an all-to-all fabric built on NeuronLink-v4. AWS publishes up to 362 FP8 PFLOPS, up to 20.7 TB of HBM3e and 706 TB/s of aggregate memory bandwidth, and states over twice the chip density per rack of Trn2. The figures stored here are chip count times AWS's published per-chip specifications, which is AWS's own arithmetic: all three of its published aggregates reconcile exactly against the Trainium3 chip at 2,517 MXFP8 TFLOPS, 144 GB of HBM3e and 4.9 TB/s. AWS labels dense and sparse itself; the sparse figures apply to FP16, BF16 and TF32 but explicitly not to FP8, and they are 3.75x the dense rate rather than twice it, because AWS's sparse mode runs at the chip's low-precision rate whatever format is fed in. AWS publishes no power figure for the UltraServer or for the chip, so total system power is left blank rather than estimated. One inconsistency in AWS's own material is worth noting: this machine's page gives NeuronLink-v4 as 2 TB/s per chip while the Trainium3 architecture documentation gives 2.56 TB/s per device. AWS describes only this one Trn3 UltraServer configuration; there is no 64-chip variant and no Gen1/Gen2 naming in its documentation.

We will point you at suppliers who have it. Free, and no signup.

FP32
26.35
PFLOPS
TF32
96.62
PFLOPS
FP16
96.62
PFLOPS
BF16
96.62
PFLOPS

System Details

GPU Configuration

GPU Model: AWS Trainium3
GPU Count: 144 GPUs
Architecture: NeuronCore-v4
Interconnect: NeuronSwitch-v1 all-to-all fabric over NeuronLink-v4, 2 TB/s per chip

System Specifications

Form Factor: rack
Total Power:
Total Memory: 21234 GB
Memory Bandwidth: 705600 GB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
26.352 PFLOPS 183.0 TFLOPS
96.624 PFLOPS 671.0 TFLOPS
96.624 PFLOPS 671.0 TFLOPS
96.624 PFLOPS 671.0 TFLOPS
362.448 PFLOPS 2517.0 TFLOPS
362.448 PFLOPS 2517.0 TFLOPS

All figures are dense. Vendors commonly headline the number, which is twice the dense one.

Powered by AWS Trainium3

This system utilizes 144 × AWS Trainium3 GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

Per GPU Memory

144 GB

Process Node

3nm

Architecture

NeuronCore-v4

Documentation & Resources

Official Datasheet

AWS Trn3 UltraServer technical specifications

Download PDF

Trainium2 Architecture (AWS Neuron Documentation)

AWS • Latest version

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The AWS Trn3 UltraServer runs 144× Trainium3 GPUs over NeuronSwitch-v1 all-to-all fabric over NeuronLink-v4, 2 TB/s per chip, delivering 362 PFLOPS FP8 dense.