AWS logo

AWS Trn3 UltraServer

rack
144× Trainium3
NeuronSwitch-v1 all-to-all fabric over NeuronLink-v4, 2 TB/s per chip
2025
FP32
26.4
PFLOPS
TF32
96.6
PFLOPS
Power
—
kW Total
Memory
20736
GB Total

The Trn3 UltraServer is AWS's rack-scale Trainium3 machine, generally available since 2 December 2025 and reaching 144 chips joined by NeuronSwitch-v1, an all-to-all fabric built on NeuronLink-v4. AWS publishes up to 362 FP8 PFLOPS, up to 20.7 TB of HBM3e and 706 TB/s of aggregate memory bandwidth, and states over twice the chip density per rack of Trn2. The figures stored here are chip count times AWS's published per-chip specifications, which is AWS's own arithmetic: all three of its published aggregates reconcile exactly against the Trainium3 chip at 2,517 MXFP8 TFLOPS, 144 GB of HBM3e and 4.9 TB/s. AWS labels dense and sparse itself; the sparse figures apply to FP16, BF16 and TF32 but explicitly not to FP8, and they are 3.75x the dense rate rather than twice it, because AWS's sparse mode runs at the chip's low-precision rate whatever format is fed in. AWS publishes no power figure for the UltraServer or for the chip, so total system power is left blank rather than estimated. One inconsistency in AWS's own material is worth noting: this machine's page gives NeuronLink-v4 as 2 TB/s per chip while the Trainium3 architecture documentation gives 2.56 TB/s per device. AWS describes only this one Trn3 UltraServer configuration; there is no 64-chip variant and no Gen1/Gen2 naming in its documentation.

We will point you at suppliers who have it. Free, and no signup.

FP32
26.4
PFLOPS
TF32
96.6
PFLOPS
FP16
96.6
PFLOPS
BF16
96.6
PFLOPS

System Details

GPU Configuration

GPU Model: AWS Trainium3
GPU Count: 144 GPUs
Architecture: NeuronCore-v4
Interconnect: NeuronSwitch-v1 all-to-all fabric over NeuronLink-v4, 2 TB/s per chip

System Specifications

Form Factor: rack
Total Power: —
Total Memory: 20.7 TB
Memory Bandwidth: 705.6 TB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
26.4 PFLOPS 183 TFLOPS —
96.6 PFLOPS 671 TFLOPS —
96.6 PFLOPS 671 TFLOPS —
96.6 PFLOPS 671 TFLOPS —
362 PFLOPS 2,517 TFLOPS —
362 PFLOPS 2,517 TFLOPS —

Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.

Powered by AWS Trainium3

This system utilizes 144 × AWS Trainium3 GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

—

Per GPU Memory

144 GB

Process Node

3nm

Architecture

NeuronCore-v4

Documentation & Resources

Flopper Spec Sheet

AWS Trn3 UltraServer specifications, generated from our database. Printable.

Vendor product page

AWS Trn3 UltraServer documentation

View ↗

AWS announces Amazon EC2 Inf2 instances (Preview)

AWS • 2022-11-29

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The AWS Trn3 UltraServer runs 144× Trainium3 GPUs over NeuronSwitch-v1 all-to-all fabric over NeuronLink-v4, 2 TB/s per chip, delivering 362 PFLOPS FP8 dense.

© 2026 Flopper.io - Compare the GPUs Powering AI