Sold in mainland China. Huawei is on the US Entity List and Ascend silicon is export-controlled wherever it is located, so this system cannot be rented from any provider on this site.
Huawei logo

Huawei Atlas 900 A3 SuperPoD

pod
384× Ascend 910C
UnifiedBus (UB) all-to-all optical; 784 GB/s bidirectional D2D per NPU
2025
FP16
307
PFLOPS
BF16
307
PFLOPS
Power
—
kW Total
Memory
49152
GB Total

Huawei's 384-NPU supernode, built from twelve 47U liquid-cooled compute cabinets plus four air-cooled bus-equipment cabinets, each 2,250 by 600 by 1,150 mm. It carries up to 384 Ascend 910 NPUs and 192 Kunpeng 920 CPUs, with 384 x 128 GB of on-chip memory at up to 3.2 TB/s each and 784 GB/s of bidirectional device-to-device bandwidth. Huawei rates it at up to 307.2 PFLOPS of FP16, and also lists 288.7 and 240.3 PFLOPS variants; the top figure works out at exactly 800 TFLOPS per NPU, which is the Ascend 910C's own published rate, so the BF16 and INT8 figures here are that chip's published 800 and 1,600 TFLOPS multiplied by 384. Cooling is liquid for the compute cabinets and air for the bus cabinets, with an operating range of 5 to 40 C. Note that Huawei publishes no power figure for this machine: its supply row states only the input voltage, three-phase 380V AC on two feeds. A figure of 559 kW previously held here appeared in no Huawei source and has been removed rather than shipped. Huawei labels nothing dense or sparse. This machine is also sold as CloudMatrix 384, which is a deployment name for the same 384-NPU configuration rather than a separate product.

We will point you at suppliers who have it. Free, and no signup.

FP16
307
PFLOPS
BF16
307
PFLOPS
INT8
614
POPS

System Details

GPU Configuration

GPU Count: 384 GPUs
Architecture: Da Vinci v2 (dual-die)
Interconnect: UnifiedBus (UB) all-to-all optical; 784 GB/s bidirectional D2D per NPU

System Specifications

Form Factor: pod
Total Power: —
Total Memory: 49.2 TB
Memory Bandwidth: 1.2 PB/s

Precision Performance Breakdown

PrecisionSystem PerformancePer GPUEfficiency
307 PFLOPS 800 TFLOPS —
307 PFLOPS 800 TFLOPS —
614 POPS 1,600 TOPS —

Where a vendor states a basis, the figure shown is the dense one, and any figure it publishes is listed separately. Vendors commonly headline the sparse number instead: for NVIDIA's 2:4 structured sparsity that is exactly twice the dense figure, though other vendors' sparse modes do not all follow that ratio.

Powered by Huawei Ascend 910C

This system utilizes 384 × Huawei Ascend 910C GPUs, each delivering exceptional performance for AI and HPC workloads.

Per GPU TDP

600W

Per GPU Memory

128 GB

Process Node

SMIC 7nm

Architecture

Da Vinci v2 (dual-die)

Documentation & Resources

Flopper Spec Sheet

Huawei Atlas 900 A3 SuperPoD specifications, generated from our database. Printable.

Vendor product page

Huawei Atlas 900 A3 SuperPoD documentation

View ↗

Huawei HiSilicon Da Vinci architecture, Hot Chips 31 (2019)

Huawei • 2019-08-19

View Document ↗

Typical Use Cases

AI/ML Training
High-Performance Computing
Data Analytics

The Huawei Atlas 900 A3 SuperPoD runs 384× Ascend 910C GPUs over UnifiedBus (UB) all-to-all optical; 784 GB/s bidirectional D2D per NPU, delivering 307 PFLOPS BF16 dense.

© 2026 Flopper.io - Compare the hardware powering AI