HW

Huawei Atlas 800I A3

rack
8× Ascend 910C
UnifiedBus (UB): 784 GB/s bidirectional D2D per NPU; 8 x 400GE QSFP-DD direct from NPU over RoCE, plus 56 x 400GE QSFP-DD over the bus protocol
2025
FP16
4.5
PFLOPS
Power
14.6
kW Total
Memory
1049
GB Total

The inference member of Huawei's Atlas 800 A3 family: a 10U air-cooled rack server with eight Ascend 910 NPUs and four Kunpeng 920 CPUs, rated at 4.48 PFLOPS of FP16 with 1,024 GB of on-chip memory across the eight NPUs at 3.2 TB/s each. That works out at 560 TFLOPS per NPU against the 750 of the training-oriented Atlas 800T A3 in the same chassis, which is a genuine inference bin rather than a transcription difference: Huawei's own family page lists three Atlas 800 A3 models at 6.0, 5.0 and 4.48 PFLOPS FP16. Memory is offered as 8 x 128 GB or 8 x 64 GB; the 128 GB configuration is recorded here. Networking is eight 400GE QSFP-DD ports direct from the NPUs over RoCE plus fifty-six more over Huawei's bus protocol, with 784 GB/s of bidirectional device-to-device bandwidth. Power is up to 14.6 kW from six hot-swap 3.0 kW supplies in 5+1 redundancy, air cooled. Huawei publishes no INT8 figure for this machine and labels nothing dense or sparse, so only the FP16 rate it states is recorded and nothing is derived.

We will point you at suppliers who have it. Free, and no signup.

FP16
4.48
PFLOPS

System Details

GPU Configuration

GPU Count: 8 GPUs
Architecture: Da Vinci v2 (dual-die)
Interconnect: UnifiedBus (UB): 784 GB/s bidirectional D2D per NPU; 8 x 400GE QSFP-DD direct from NPU over RoCE, plus 56 x 400GE QSFP-DD over the bus protocol

System Specifications

Form Factor: rack
Total Power: 14.6 kW
Total Memory: 1049 GB
Memory Bandwidth: 25600 GB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
4.480 PFLOPS 560.0 TFLOPS

All figures are dense. Vendors commonly headline the number, which is twice the dense one.

Powered by Huawei Ascend 910C

This system utilizes 8 × Huawei Ascend 910C GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

600W

Per GPU Memory

128 GB

Process Node

SMIC 7nm

Architecture

Da Vinci v2 (dual-die)

Documentation & Resources

Official Datasheet

Huawei Atlas 800I A3 technical specifications

Download PDF

Atlas 650E 服务器 技术规格 (Huawei Ascend AI server product page)

Huawei • Latest version

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The Huawei Atlas 800I A3 runs 8× Ascend 910C GPUs over UnifiedBus (UB): 784 GB/s bidirectional D2D per NPU; 8 x 400GE QSFP-DD direct from NPU over RoCE, plus 56 x 400GE QSFP-DD over the bus protocol, delivering 4.48 PFLOPS FP16 dense.