NVIDIA A100 PCIe 40GB vs NVIDIA A800 SXM4

The A100 and the A800 deliver near-identical FP16 throughput (312 vs 312 TFLOPS dense). The A800 also carries 40 GB more memory (80 GB vs 40 GB).

A100: Ampere, 2020 A800: Ampere, 2022
Memory
40 vs 80 GB
40 GB more for the A800
Bandwidth
1.6 TB/s vs 2.0 TB/s
A800 moves data faster
TDP
250 W vs 400 W
A100 draws 150 W less
See rental pricing from $0.53/hr Download comparison sheet

Share this page

NVIDIA A100 PCIe 40GB vs NVIDIA A800 SXM4

A post with this link shows the card on the Image tab as its preview.

Printable version →

Specifications

A100
A800
Architecture
AmpereAmpere
Launch Year
20202022
Form Factor
PCIeSXM
Memory
40 GB80 GB
Memory Bandwidth
1.6 TB/s2.0 TB/s
TDP
250 W400 W
Process Node
7nm7nm

Performance (TFLOPS)

A100
A800
FP64
9.7 TFLOPS 9.7 TFLOPS
FP32
20 TFLOPS 20 TFLOPS
TF32
156 TFLOPS
312 TFLOPS sparse
156 TFLOPS
312 TFLOPS sparse
BF16
312 TFLOPS
624 TFLOPS sparse
312 TFLOPS
624 TFLOPS sparse
FP16
312 TFLOPS
624 TFLOPS sparse
312 TFLOPS
624 TFLOPS sparse
FP8
No verified data available No verified data available
FP6
No verified data available No verified data available
FP4
No verified data available No verified data available
INT8
624 TOPS
1,248 TOPS sparse
624 TOPS
1,248 TOPS sparse

FLOPS by Precision

What actually differs

The A800 is the newer part: Ampere, launched in 2022, against the A100's Ampere from 2020. Newer architectures typically add lower-precision formats and better throughput per watt, so check the precision rows your workload actually uses.

Dense throughput for the A100 against the A800: FP64 9.7 vs 9.7 TFLOPS, FP32 20 vs 20 TFLOPS, FP16 312 vs 312 TFLOPS. Sparse figures, where the vendor publishes them, appear under each dense number in the performance table.

Memory is 40 GB against 80 GB, fed at 1.6 TB/s versus 2.0 TB/s. More memory per GPU means larger models fit before you have to shard across cards, which often matters more than raw TFLOPS for inference.

Power budgets are 250 W for the A100 and 400 W for the A800. At FP32 that works out to 0.08 against 0.05 TFLOPS per watt.

For LLM training and inference, weight the FP8 and FP16 rows and memory capacity most heavily. For scientific and HPC workloads, the FP64 row is the one to read.

Frequently asked questions

Is the A100 faster than the A800?

They are close on paper: both deliver about 312 TFLOPS of dense FP16 throughput. Memory, bandwidth, and power are the deciding differences.

Which has more memory, the A100 or the A800?

The A800 carries 80 GB of memory versus 40 GB for the A100. Memory bandwidth is 1.6 TB/s for the A100 and 2.0 TB/s for the A800.

How much power do the A100 and the A800 draw?

The A100 is rated at 250 W TDP and the A800 at 400 W. On FP32 throughput per watt, the A100 is the more efficient part.

Can I rent the A100 or the A800 in the cloud?

Yes. Live cloud listings tracked by Flopper start at $0.53 per GPU hour. The rental pricing section on this page lists current providers and rates for both GPUs.

Where to Rent

NVIDIA A100 PCIe 40GB

ProviderConfigurationPrice/GPU-hrChecked
4× A100 PCIe Community $0.53 1h ago View →
All A100 listings and price history →

Get Comparison Updates

New GPUs added weekly. Be the first to see how they compare.

© 2026 Flopper.io - Compare the GPUs Powering AI