Huawei Atlas 650E
A 14U air-cooled rack server with eight Ascend 950DT NPUs and two Kunpeng 950 CPUs, built for ordinary air-cooled machine rooms rather than liquid-cooled halls. Huawei rates it at 12.48 PFLOPS mxFP4, 6.43 PFLOPS across mxFP8, FP8 and HiF8, and 3.40 PFLOPS FP16/BF16, with 768 GB of on-chip HBM across the eight NPUs at up to 4.0 TB/s each. Two machines can be joined over UnifiedBus into a sixteen-NPU full-mesh. Power is up to 14.5 kW from six hot-swap 3.0 kW supplies in 5+1 redundancy, cooled by thirty hot-swap fan modules. Note that the per-NPU figure works out at 1.56 PFLOPS mxFP4, below the 1.95 PFLOPS the same Ascend 950DT reaches in the liquid-cooled Atlas 950 SuperPoD: this is a de-rated air-cooled bin, not a transcription error, and should not be reconciled upward. Huawei labels none of these figures dense or sparse and publishes no sparse rate; its asterisk on the compute row means only that the values are theoretical and measured results may differ by under one percent.
We will point you at suppliers who have it. Free, and no signup.
System Details
GPU Configuration
System Specifications
Precision Performance Breakdown
| Precision | System Performance | Per GPU | Efficiency |
|---|---|---|---|
| 3.400 PFLOPS | 425.0 TFLOPS | — | |
| 3.400 PFLOPS | 425.0 TFLOPS | — | |
| 6.430 PFLOPS | 803.8 TFLOPS | — | |
| 12.480 PFLOPS | 1560.0 TFLOPS | — |
All figures are dense. Vendors commonly headline the number, which is twice the dense one.
Powered by Huawei Ascend 950DT
This system utilizes 8 × Huawei Ascend 950DT GPUs, each delivering exceptional performance for AI and HPC workloads.
Per GPU TDP
—
Per GPU Memory
144 GB
Process Node
—
Architecture
Da Vinci v3
Related Systems
Documentation & Resources
Official Datasheet
Huawei Atlas 650E technical specifications
Atlas 650E 服务器 技术规格 (Huawei Ascend AI server product page)
Huawei • Latest version
Typical Use Cases
The Huawei Atlas 650E runs 8× Ascend 950DT GPUs over UnifiedBus (Lingqu): 8 x 784 GB/s bidirectional within a node, 2 x 8 x 1.68 TB/s across a two-node pair; UBoE direct from NPU and RoCE both at 400 Gbps per NPU, delivering 6.43 PFLOPS FP8 dense.