NVIDIA A16 16GB
Overview
The NVIDIA A16 is a quad-GPU board built for graphics-rich virtual desktops, and it is in this database because it turns up on GPU marketplaces as four independently rentable accelerators rather than as one card. Each of its four GA107 GPUs carries 1,280 Ampere CUDA cores, 40 third-generation Tensor Cores and 16 GB of GDDR6 at 200 GB/s, which NVIDIA rates at 4.5 TFLOPS FP32, 9 TFLOPS dense TF32, 17.9 TFLOPS dense FP16 and 35.9 dense INT8 TOPS. That is a modest accelerator by datacenter standards and it is not trying to be anything else: the A16 exists to put a real GPU and a decent frame buffer in front of as many concurrent users as possible, up to 64 per board, with four hardware encoders and eight decoders to feed them. For inference work the useful comparison is the ratio rather than the absolute figure, since its Tensor Cores sit in the professional Ampere bucket, running TF32 at twice the shader rate and FP16 at four times it, with structured sparsity doubling both again. The whole board draws 250 W, so four of these GPUs cost about what a single mid-range datacenter card does.
Performance
Peak theoretical throughput by precision type
| Precision | Dense | 2:4 Sparse |
|---|---|---|
FP64 | No verified data available | Structured sparsity is a tensor-core feature; this vector precision has no sparse form |
FP32 32-bit floating point | 4.5TFLOPS | Structured sparsity is a tensor-core feature; this vector precision has no sparse form |
TF32 TensorFloat-32 | 9TFLOPS | 18TFLOPS |
BF16 | No verified data available | The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision |
FP16 16-bit floating point | 18TFLOPS | 36TFLOPS |
FP8 | No verified data available | The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision |
FP6 | No verified data available | The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision |
FP4 | No verified data available | The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision |
INT8 8-bit integer | 36TOPS | 72TOPS |
Specifications
Architecture
Ampere
Form Factor
PCIe
Launch Year
2021
Process Node
No verified data available
Memory
16 GB GDDR6
Bandwidth
200 GB/s
TDP
No verified data available
Max power
No verified data available
CUDA Cores
1,280
Spec Confidence
Official
Full Specifications
| Compute Engine | |
|---|---|
| CUDA Cores | 1,280 |
| Tensor Cores | 40 (3rd Gen) |
| Streaming Multiprocessors | 10 |
| Memory | |
| Memory | 16 GB |
| Memory Type | GDDR6 |
| Bandwidth | 200 GB/s |
| Interface Width | No verified data available |
| Interconnect & I/O | |
| PCIe | 4.0 x16 |
| Power & Thermal | |
| TDP | No verified data available |
| Max power | No verified data available |
| Cooling | Passive |
| Enterprise Features | |
| ECC Memory | Yes |
| MIG Support | No |
| Sparsity | Yes |
| Compute APIs | CUDA, DirectCompute, OpenCL, OpenACC |
| Physical & Media | |
| NVENC Engines | 1 |
| NVDEC Engines | 2 |
| General | |
| Form Factor | PCIe |
| Architecture | Ampere |
| Process Node | No verified data available |
| Launch Year | 2021 |
Datasheet & Resources
Data Provenance
Every figure traced to a source
Primary Source
- Document
- NVIDIA A16 datasheet
- Publisher
- NVIDIA
- Published
- 2022-02-09
Data Quality
- Spec confidence
- Official
- Clock basis
- Boost
- Core precisions with figures
- 4 of 9
- Normalization
- All values in TFLOPS
Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.
Browse GPU Cloud ProvidersSimilar GPUs
Frequently Asked Questions
How many TFLOPS does the NVIDIA A16 have?
The NVIDIA A16 delivers 4.5 TFLOPS FP32 and 18 TFLOPS FP16 at peak. Flopper does not currently have a verified FP8 throughput figure for it.
What is the power consumption of the NVIDIA A16?
Flopper does not currently have a verified TDP (Thermal Design Power) figure for the NVIDIA A16.
How much memory does the NVIDIA A16 have?
The NVIDIA A16 is equipped with 16 GB of memory with 200 GB/s of memory bandwidth.
What architecture is the NVIDIA A16 based on?
The NVIDIA A16 is based on the Ampere architecture, launched in 2021.
Stay Updated on GPU Releases
Get notified when new GPUs are added or specifications are updated.
No spam, unsubscribe anytime.