Sold in mainland China. Huawei is on the US Entity List and Ascend silicon is export-controlled wherever it is located, so this server cannot be rented from any provider on this site.
Huawei logo

Huawei Atlas 650E

rack
8× Ascend 950DT
UnifiedBus (Lingqu): 8 x 784 GB/s bidirectional within a node, 2 x 8 x 1.68 TB/s across a two-node pair; UBoE direct from NPU and RoCE both at 400 Gbps per NPU
2026
FP16
3,400
TFLOPS
BF16
3,400
TFLOPS
Power
14.5
kW Total
Memory
768
GB Total

A 14U air-cooled rack server with eight Ascend 950DT NPUs and two Kunpeng 950 CPUs, built for ordinary air-cooled machine rooms rather than liquid-cooled halls. Huawei rates it at 12.48 PFLOPS mxFP4, 6.43 PFLOPS across mxFP8, FP8 and HiF8, and 3.40 PFLOPS FP16/BF16, with 768 GB of on-chip HBM across the eight NPUs at up to 4.0 TB/s each. Two machines can be joined over UnifiedBus into a sixteen-NPU full-mesh. Power is up to 14.5 kW from six hot-swap 3.0 kW supplies in 5+1 redundancy, cooled by thirty hot-swap fan modules. Note that the per-NPU figure works out at 1.56 PFLOPS mxFP4, below the 1.95 PFLOPS the same Ascend 950DT reaches in the liquid-cooled Atlas 950 SuperPoD: this is a de-rated air-cooled bin, not a transcription error, and should not be reconciled upward. Huawei labels none of these figures dense or sparse and publishes no sparse rate; its asterisk on the compute row means only that the values are theoretical and measured results may differ by under one percent.

We will point you at suppliers who have it. Free, and no signup.

FP16
3,400
TFLOPS
BF16
3,400
TFLOPS
FP8
6,430
TFLOPS
FP4
12,480
TFLOPS

System Details

GPU Configuration

GPU Count: 8 GPUs
Architecture: Da Vinci v3
Interconnect: UnifiedBus (Lingqu): 8 x 784 GB/s bidirectional within a node, 2 x 8 x 1.68 TB/s across a two-node pair; UBoE direct from NPU and RoCE both at 400 Gbps per NPU

System Specifications

Form Factor: rack
Total Power: 14.5 kW
Total Memory: 768 GB
Memory Bandwidth: 32 TB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
3,400 TFLOPS 425 TFLOPS —
3,400 TFLOPS 425 TFLOPS —
6,430 TFLOPS 804 TFLOPS —
12,480 TFLOPS 1,560 TFLOPS —

Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.

Powered by Huawei Ascend 950DT

This system utilizes 8 × Huawei Ascend 950DT GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

—

Per GPU Memory

144 GB

Process Node

—

Architecture

Da Vinci v3

Documentation & Resources

Flopper Spec Sheet

Huawei Atlas 650E specifications, generated from our database. Printable.

Vendor product page

Huawei Atlas 650E documentation

View ↗

Huawei HiSilicon Da Vinci architecture, Hot Chips 31 (2019)

Huawei • 2019-08-19

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The Huawei Atlas 650E runs 8× Ascend 950DT GPUs over UnifiedBus (Lingqu): 8 x 784 GB/s bidirectional within a node, 2 x 8 x 1.68 TB/s across a two-node pair; UBoE direct from NPU and RoCE both at 400 Gbps per NPU, delivering 6,430 TFLOPS FP8 dense.

© 2026 Flopper.io - Compare the hardware powering AI