AWS Trn3 UltraServer
The Trn3 UltraServer is AWS's rack-scale Trainium3 machine, generally available since 2 December 2025 and reaching 144 chips joined by NeuronSwitch-v1, an all-to-all fabric built on NeuronLink-v4. AWS publishes up to 362 FP8 PFLOPS, up to 20.7 TB of HBM3e and 706 TB/s of aggregate memory bandwidth, and states over twice the chip density per rack of Trn2. The figures stored here are chip count times AWS's published per-chip specifications, which is AWS's own arithmetic: all three of its published aggregates reconcile exactly against the Trainium3 chip at 2,517 MXFP8 TFLOPS, 144 GB of HBM3e and 4.9 TB/s. AWS labels dense and sparse itself; the sparse figures apply to FP16, BF16 and TF32 but explicitly not to FP8, and they are 3.75x the dense rate rather than twice it, because AWS's sparse mode runs at the chip's low-precision rate whatever format is fed in. AWS publishes no power figure for the UltraServer or for the chip, so total system power is left blank rather than estimated. One inconsistency in AWS's own material is worth noting: this machine's page gives NeuronLink-v4 as 2 TB/s per chip while the Trainium3 architecture documentation gives 2.56 TB/s per device. AWS describes only this one Trn3 UltraServer configuration; there is no 64-chip variant and no Gen1/Gen2 naming in its documentation.
We will point you at suppliers who have it. Free, and no signup.
System Details
GPU Configuration
System Specifications
Precision Performance Breakdown
| Precision | System Performance | Per GPU | Efficiency |
|---|---|---|---|
| 26.352 PFLOPS | 183.0 TFLOPS | — | |
| 96.624 PFLOPS | 671.0 TFLOPS | — | |
| 96.624 PFLOPS | 671.0 TFLOPS | — | |
| 96.624 PFLOPS | 671.0 TFLOPS | — | |
| 362.448 PFLOPS | 2517.0 TFLOPS | — | |
| 362.448 PFLOPS | 2517.0 TFLOPS | — |
All figures are dense. Vendors commonly headline the number, which is twice the dense one.
Powered by AWS Trainium3
This system utilizes 144 × AWS Trainium3 GPUs, each delivering exceptional performance for AI and HPC workloads.
Per GPU TDP
—
Per GPU Memory
144 GB
Process Node
3nm
Architecture
NeuronCore-v4
Related Systems
Documentation & Resources
Official Datasheet
AWS Trn3 UltraServer technical specifications
Trainium2 Architecture (AWS Neuron Documentation)
AWS • Latest version
Typical Use Cases
The AWS Trn3 UltraServer runs 144× Trainium3 GPUs over NeuronSwitch-v1 all-to-all fabric over NeuronLink-v4, 2 TB/s per chip, delivering 362 PFLOPS FP8 dense.