
AWS Trn3 UltraServer
The Trn3 UltraServer is AWS's rack-scale Trainium3 machine, generally available since 2 December 2025 and reaching 144 chips joined by NeuronSwitch-v1, an all-to-all fabric built on NeuronLink-v4. AWS publishes up to 362 FP8 PFLOPS, up to 20.7 TB of HBM3e and 706 TB/s of aggregate memory bandwidth, and states over twice the chip density per rack of Trn2. The figures stored here are chip count times AWS's published per-chip specifications, which is AWS's own arithmetic: all three of its published aggregates reconcile exactly against the Trainium3 chip at 2,517 MXFP8 TFLOPS, 144 GB of HBM3e and 4.9 TB/s. AWS labels dense and sparse itself; the sparse figures apply to FP16, BF16 and TF32 but explicitly not to FP8, and they are 3.75x the dense rate rather than twice it, because AWS's sparse mode runs at the chip's low-precision rate whatever format is fed in. AWS publishes no power figure for the UltraServer or for the chip, so total system power is left blank rather than estimated. One inconsistency in AWS's own material is worth noting: this machine's page gives NeuronLink-v4 as 2 TB/s per chip while the Trainium3 architecture documentation gives 2.56 TB/s per device. AWS describes only this one Trn3 UltraServer configuration; there is no 64-chip variant and no Gen1/Gen2 naming in its documentation.
We will point you at suppliers who have it. Free, and no signup.
System Details
GPU Configuration
System Specifications
Precision Performance Breakdown
| Precision | System Performance | Per GPU | Efficiency |
|---|---|---|---|
| 26.4 PFLOPS | 183 TFLOPS | — | |
| 96.6 PFLOPS | 671 TFLOPS | — | |
| 96.6 PFLOPS | 671 TFLOPS | — | |
| 96.6 PFLOPS | 671 TFLOPS | — | |
| 362 PFLOPS | 2,517 TFLOPS | — | |
| 362 PFLOPS | 2,517 TFLOPS | — |
Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.
Powered by AWS Trainium3
This system utilizes 144 × AWS Trainium3 GPUs, each delivering exceptional performance for AI and HPC workloads.
Per GPU TDP
—
Per GPU Memory
144 GB
Process Node
3nm
Architecture
NeuronCore-v4
Related Systems
Documentation & Resources
Flopper Spec Sheet
AWS Trn3 UltraServer specifications, generated from our database. Printable.
Vendor product page
AWS Trn3 UltraServer documentation
AWS announces Amazon EC2 Inf2 instances (Preview)
AWS • 2022-11-29
Typical Use Cases
The AWS Trn3 UltraServer runs 144× Trainium3 GPUs over NeuronSwitch-v1 all-to-all fabric over NeuronLink-v4, 2 TB/s per chip, delivering 362 PFLOPS FP8 dense.