HW

Huawei Atlas 650E

rack
8× Ascend 950DT
UnifiedBus (Lingqu): 8 x 784 GB/s bidirectional within a node, 2 x 8 x 1.68 TB/s across a two-node pair; UBoE direct from NPU and RoCE both at 400 Gbps per NPU
2026
FP16
3.4
PFLOPS
BF16
3.4
PFLOPS
Power
14.5
kW Total
Memory
786
GB Total

A 14U air-cooled rack server with eight Ascend 950DT NPUs and two Kunpeng 950 CPUs, built for ordinary air-cooled machine rooms rather than liquid-cooled halls. Huawei rates it at 12.48 PFLOPS mxFP4, 6.43 PFLOPS across mxFP8, FP8 and HiF8, and 3.40 PFLOPS FP16/BF16, with 768 GB of on-chip HBM across the eight NPUs at up to 4.0 TB/s each. Two machines can be joined over UnifiedBus into a sixteen-NPU full-mesh. Power is up to 14.5 kW from six hot-swap 3.0 kW supplies in 5+1 redundancy, cooled by thirty hot-swap fan modules. Note that the per-NPU figure works out at 1.56 PFLOPS mxFP4, below the 1.95 PFLOPS the same Ascend 950DT reaches in the liquid-cooled Atlas 950 SuperPoD: this is a de-rated air-cooled bin, not a transcription error, and should not be reconciled upward. Huawei labels none of these figures dense or sparse and publishes no sparse rate; its asterisk on the compute row means only that the values are theoretical and measured results may differ by under one percent.

We will point you at suppliers who have it. Free, and no signup.

FP16
3.40
PFLOPS
BF16
3.40
PFLOPS
FP8
6.43
PFLOPS
FP4
12.48
PFLOPS

System Details

GPU Configuration

GPU Count: 8 GPUs
Architecture: Da Vinci v3
Interconnect: UnifiedBus (Lingqu): 8 x 784 GB/s bidirectional within a node, 2 x 8 x 1.68 TB/s across a two-node pair; UBoE direct from NPU and RoCE both at 400 Gbps per NPU

System Specifications

Form Factor: rack
Total Power: 14.5 kW
Total Memory: 786 GB
Memory Bandwidth: 32000 GB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
3.400 PFLOPS 425.0 TFLOPS
3.400 PFLOPS 425.0 TFLOPS
6.430 PFLOPS 803.8 TFLOPS
12.480 PFLOPS 1560.0 TFLOPS

All figures are dense. Vendors commonly headline the number, which is twice the dense one.

Powered by Huawei Ascend 950DT

This system utilizes 8 × Huawei Ascend 950DT GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

Per GPU Memory

144 GB

Process Node

Architecture

Da Vinci v3

Documentation & Resources

Official Datasheet

Huawei Atlas 650E technical specifications

Download PDF

Atlas 650E 服务器 技术规格 (Huawei Ascend AI server product page)

Huawei • Latest version

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The Huawei Atlas 650E runs 8× Ascend 950DT GPUs over UnifiedBus (Lingqu): 8 x 784 GB/s bidirectional within a node, 2 x 8 x 1.68 TB/s across a two-node pair; UBoE direct from NPU and RoCE both at 400 Gbps per NPU, delivering 6.43 PFLOPS FP8 dense.