
Huawei Atlas 650E
A 14U air-cooled rack server with eight Ascend 950DT NPUs and two Kunpeng 950 CPUs, built for ordinary air-cooled machine rooms rather than liquid-cooled halls. Huawei rates it at 12.48 PFLOPS mxFP4, 6.43 PFLOPS across mxFP8, FP8 and HiF8, and 3.40 PFLOPS FP16/BF16, with 768 GB of on-chip HBM across the eight NPUs at up to 4.0 TB/s each. Two machines can be joined over UnifiedBus into a sixteen-NPU full-mesh. Power is up to 14.5 kW from six hot-swap 3.0 kW supplies in 5+1 redundancy, cooled by thirty hot-swap fan modules. Note that the per-NPU figure works out at 1.56 PFLOPS mxFP4, below the 1.95 PFLOPS the same Ascend 950DT reaches in the liquid-cooled Atlas 950 SuperPoD: this is a de-rated air-cooled bin, not a transcription error, and should not be reconciled upward. Huawei labels none of these figures dense or sparse and publishes no sparse rate; its asterisk on the compute row means only that the values are theoretical and measured results may differ by under one percent.
We will point you at suppliers who have it. Free, and no signup.
System Details
GPU Configuration
System Specifications
Precision Performance Breakdown
| Precision | System Performance | Per GPU | Efficiency |
|---|---|---|---|
| 3,400 TFLOPS | 425 TFLOPS | — | |
| 3,400 TFLOPS | 425 TFLOPS | — | |
| 6,430 TFLOPS | 804 TFLOPS | — | |
| 12,480 TFLOPS | 1,560 TFLOPS | — |
Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.
Powered by Huawei Ascend 950DT
This system utilizes 8 × Huawei Ascend 950DT GPUs, each delivering exceptional performance for AI and HPC workloads.
Per GPU TDP
—
Per GPU Memory
144 GB
Process Node
—
Architecture
Da Vinci v3
Documentation & Resources
Flopper Spec Sheet
Huawei Atlas 650E specifications, generated from our database. Printable.
Vendor product page
Huawei Atlas 650E documentation
Huawei HiSilicon Da Vinci architecture, Hot Chips 31 (2019)
Huawei • 2019-08-19
Typical Use Cases
The Huawei Atlas 650E runs 8× Ascend 950DT GPUs over UnifiedBus (Lingqu): 8 x 784 GB/s bidirectional within a node, 2 x 8 x 1.68 TB/s across a two-node pair; UBoE direct from NPU and RoCE both at 400 Gbps per NPU, delivering 6,430 TFLOPS FP8 dense.