Sold in mainland China. Huawei is on the US Entity List and Ascend silicon is export-controlled wherever it is located, so this system cannot be rented from any provider on this site.
Huawei logo

Huawei Atlas 800T A3

rack
8× Ascend 910C
UnifiedBus (UB): 784 GB/s bidirectional D2D per NPU; scales to a 384-NPU Atlas 900 A3 SuperPoD
2025
FP16
6,000
TFLOPS
Power
—
kW Total
Memory
1024
GB Total

The training member of Huawei's Atlas 800 A3 family: a 10U air-cooled compute node with eight Ascend 910 NPUs and four Kunpeng 920 CPUs, rated at 6.0 PFLOPS of FP16 with 1,024 GB of on-chip memory at 3.2 TB/s per NPU and 784 GB/s of bidirectional device-to-device bandwidth. Huawei sells three models on this chassis at 6.0, 5.0 and 4.48 PFLOPS FP16, which works out at 750, 625 and 560 TFLOPS per NPU; this is the top one, and the 4.48 model is the separately listed Atlas 800I A3. Multiple nodes combine into an Atlas 900 A3 SuperPoD of up to 384 cards. Huawei publishes only an FP16 figure for this machine, and no INT8: a previously held INT8 value of 12.0 PFLOPS was 8 x 1,500 TFLOPS, a per-NPU rate Huawei states nowhere, and has been removed. Power is also not attributable: Huawei's family page gives 16.2 kW as the maximum input power for the Atlas 800 A3 line, but that line covers three compute variants and the Atlas 800I A3's own page separately says 14.6 kW, so no figure is recorded here rather than guessing which model the 16.2 kW describes. Huawei labels nothing dense or sparse.

We will point you at suppliers who have it. Free, and no signup.

FP16
6,000
TFLOPS

System Details

GPU Configuration

GPU Count: 8 GPUs
Architecture: Da Vinci v2 (dual-die)
Interconnect: UnifiedBus (UB): 784 GB/s bidirectional D2D per NPU; scales to a 384-NPU Atlas 900 A3 SuperPoD

System Specifications

Form Factor: rack
Total Power: —
Total Memory: 1.0 TB
Memory Bandwidth: 25.6 TB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
6,000 TFLOPS 750 TFLOPS —

Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.

Powered by Huawei Ascend 910C

This system utilizes 8 × Huawei Ascend 910C GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

600W

Per GPU Memory

128 GB

Process Node

SMIC 7nm

Architecture

Da Vinci v2 (dual-die)

Documentation & Resources

Flopper Spec Sheet

Huawei Atlas 800T A3 specifications, generated from our database. Printable.

Vendor product page

Huawei Atlas 800T A3 documentation

View ↗

Huawei HiSilicon Da Vinci architecture, Hot Chips 31 (2019)

Huawei • 2019-08-19

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The Huawei Atlas 800T A3 runs 8× Ascend 910C GPUs over UnifiedBus (UB): 784 GB/s bidirectional D2D per NPU; scales to a 384-NPU Atlas 900 A3 SuperPoD, delivering 6,000 TFLOPS FP16 dense.

© 2026 Flopper.io - Compare the hardware powering AI