NVIDIA A800 SXM4
Overview
The NVIDIA A800 SXM4 is the mezzanine form of the A800, the export-compliance version of the A100 that NVIDIA sold into China, Hong Kong and Macau from late 2022. It is the same GA100 silicon in the same configuration as an A100 80GB SXM, and the compute figures are identical to the decimal place: 9.7 TFLOPS of FP64, 19.5 of FP32, 156 of TF32 and 312 of FP16 and BF16, with 80 GB of HBM2e at 2,039 GB/s and a 400 W board power. One thing changes, and it is the thing the export rules of the day measured. NVLink runs at 400 GB/s per GPU rather than the A100's 600 GB/s, which slows the all-to-all traffic that large-model training depends on while leaving single-GPU throughput untouched. That makes the A800 an unusually clean illustration of what an interconnect ceiling costs: identical silicon, identical arithmetic, a third less bandwidth between chips. Eight of these modules sit on an HGX A800 baseboard, where NVSwitch gives any pair of GPUs a direct 400 GB/s path, which is how they reach the market in machines such as Inspur's NF5688M6. The part had a short life. The United States tightened the rules again in October 2023, adding the A800 itself to the controlled list, and OEM product guides for it are published today as withdrawn.
Performance
Peak theoretical throughput by precision type
| Precision | Dense | 2:4 Sparse |
|---|---|---|
FP64 64-bit floating point | 9.7TFLOPS | Structured sparsity is a tensor-core feature; this vector precision has no sparse form |
FP32 32-bit floating point | 20TFLOPS | Structured sparsity is a tensor-core feature; this vector precision has no sparse form |
TF32 TensorFloat-32 | 156TFLOPS | 312TFLOPS |
BF16 Brain Float 16 | 312TFLOPS | 624TFLOPS |
FP16 16-bit floating point | 312TFLOPS | 624TFLOPS |
FP8 | No verified data available | The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision |
FP6 | No verified data available | The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision |
FP4 | No verified data available | The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision |
INT8 8-bit integer | 624TOPS | 1,248TOPS |
INT4 4-bit integer | 1,248TOPS | 2,496TOPS |
Dense peak divided by accelerator TDP. Board power only: excludes host CPUs, networking, cooling and facility overhead. TDP for this part: 400 W.
The A800 SXM4 in the GPU landscape
Peak FP16 TFLOPS (dense) against TDP, single-GPU parts tracked by Flopper
Higher and further left is better: more half-precision throughput for less power.
Specifications
Architecture
Ampere
Form Factor
SXM
Launch Year
2022
Process Node
7nm
Memory
80 GB HBM2e
Bandwidth
2,039 GB/s
TDP
400 W
Max power
500 W
Interconnect
400 GB/s NVLink
bidirectional, per GPU
CUDA Cores
6,912
Spec Confidence
Official
Full Specifications
| Compute Engine | |
|---|---|
| CUDA Cores | 6,912 |
| Tensor Cores | 432 (3rd Gen) |
| Streaming Multiprocessors | 108 |
| Memory | |
| Memory | 80 GB |
| Memory Type | HBM2e |
| Bandwidth | 2.0 TB/s |
| Interface Width | No verified data available |
| Interconnect & I/O | |
| GPU-to-GPU | NVLink |
| Interconnect Bandwidth | 400 GB/s bidirectional, per GPU |
| PCIe | Gen4 x16 |
| Power & Thermal | |
| TDP | 400 W |
| Max power | 500 W |
| Enterprise Features | |
| ECC Memory | Yes |
| MIG Support | Up to 7 MIGs @ 10GB |
| Sparsity | Yes |
| Compute APIs | CUDA, DirectCompute, OpenCL, OpenACC |
| General | |
| Form Factor | SXM |
| Architecture | Ampere |
| Process Node | 7nm |
| Launch Year | 2022 |
Datasheet & Resources
Data Provenance
Every figure traced to a source
Primary Source
- Publisher
- NVIDIA
- Published
- 2022-10-01
Data Quality
- Spec confidence
- Official
- Clock basis
- Boost
- Core precisions with figures
- 6 of 9
- Normalization
- All values in TFLOPS
Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.
Browse GPU Cloud ProvidersSimilar GPUs
Systems Using This GPU
Pre-configured systems featuring the NVIDIA A800
| System | GPU Count | Peak Performance | Total Power | |
|---|---|---|---|---|
Inspur NF5688M6 NF · rack · 2021 | 8x A800 | No verified data available | No verified data available | View System |
Frequently Asked Questions
How many TFLOPS does the NVIDIA A800 have?
The NVIDIA A800 delivers 20 TFLOPS FP32 and 312 TFLOPS FP16 at peak. Flopper does not currently have a verified FP8 throughput figure for it.
What is the power consumption of the NVIDIA A800?
The NVIDIA A800 has a TDP (Thermal Design Power) rating of 400 watts.
How much memory does the NVIDIA A800 have?
The NVIDIA A800 is equipped with 80 GB of memory with 2,039 GB/s of memory bandwidth.
What architecture is the NVIDIA A800 based on?
The NVIDIA A800 is based on the Ampere architecture, launched in 2022.
Stay Updated on GPU Releases
Get notified when new GPUs are added or specifications are updated.
No spam, unsubscribe anytime.