THIS ROW IS ONE GPU OF FOUR. The NVIDIA A16 is a single card carrying four separate GA107 GPUs, each with its own 16 GB of GDDR6, and NVIDIA writes every figure on its datasheet with a literal "4x" prefix. Everything stored here is the per-GPU figure, because that is the unit the card is rented and scheduled in: a vGPU profile, and a marketplace listing, is one of the four. POWER IS DELIBERATELY ABSENT. NVIDIA publishes exactly one power number for the A16, "Max power consumption 250W", and it covers all four dies plus the shared voltage regulation, memory and cooling of one board. Dividing it by four would publish a per-GPU figure NVIDIA has never stated, so tdp_watts and max_power_watts are NULL, following the nvidia-gb200-nvl72-gpu-186gb row. The encode and decode engine counts ARE divided, 4 NVENC and 8 NVDEC across four identical dies giving 1 and 2, because those are physically per-die. Also board-level and therefore not stored on this row: the 8-pin CPU power connector and the full-height, full-length dual-slot form factor. REFUSED: memory_interface_width and memory_clock_gbps, since NVIDIA publishes only a per-GPU bandwidth and no bus width or data rate; process_node, unpublished for GA107; base and boost clocks, unpublished.

NVIDIA A16 16GB

Ampere PCIe 2021 Spec confidence: Official
FP32
4.5
TFLOPS
Memory
16 GB
GDDR6
Bandwidth
200 GB/s
memory
Download datasheet

Overview

The NVIDIA A16 is a quad-GPU board built for graphics-rich virtual desktops, and it is in this database because it turns up on GPU marketplaces as four independently rentable accelerators rather than as one card. Each of its four GA107 GPUs carries 1,280 Ampere CUDA cores, 40 third-generation Tensor Cores and 16 GB of GDDR6 at 200 GB/s, which NVIDIA rates at 4.5 TFLOPS FP32, 9 TFLOPS dense TF32, 17.9 TFLOPS dense FP16 and 35.9 dense INT8 TOPS. That is a modest accelerator by datacenter standards and it is not trying to be anything else: the A16 exists to put a real GPU and a decent frame buffer in front of as many concurrent users as possible, up to 64 per board, with four hardware encoders and eight decoders to feed them. For inference work the useful comparison is the ratio rather than the absolute figure, since its Tensor Cores sit in the professional Ampere bucket, running TF32 at twice the shader rate and FP16 at four times it, with structured sparsity doubling both again. The whole board draws 250 W, so four of these GPUs cost about what a single mid-range datacenter card does.

Performance

Peak theoretical throughput by precision type

PrecisionDense2:4 Sparse
FP64
No verified data available Structured sparsity is a tensor-core feature; this vector precision has no sparse form
FP32
32-bit floating point
4.5TFLOPS Structured sparsity is a tensor-core feature; this vector precision has no sparse form
TF32
TensorFloat-32
9TFLOPS 18TFLOPS
BF16
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
FP16
16-bit floating point
18TFLOPS 36TFLOPS
FP8
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
FP6
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
FP4
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
INT8
8-bit integer
36TOPS 72TOPS

Specifications

Architecture

Ampere

Form Factor

PCIe

Launch Year

2021

Process Node

No verified data available

Memory

16 GB GDDR6

Bandwidth

200 GB/s

TDP

No verified data available

Max power

No verified data available

CUDA Cores

1,280

Spec Confidence

Official

Full Specifications

Compute Engine
CUDA Cores 1,280
Tensor Cores 40 (3rd Gen)
Streaming Multiprocessors 10
Memory
Memory 16 GB
Memory Type GDDR6
Bandwidth 200 GB/s
Interface Width No verified data available
Interconnect & I/O
PCIe 4.0 x16
Power & Thermal
TDP No verified data available
Max power No verified data available
Cooling Passive
Enterprise Features
ECC Memory Yes
MIG Support No
Sparsity Yes
Compute APIs CUDA, DirectCompute, OpenCL, OpenACC
Physical & Media
NVENC Engines 1
NVDEC Engines 2
General
Form Factor PCIe
Architecture Ampere
Process Node No verified data available
Launch Year 2021

Datasheet & Resources

Data Provenance

Every figure traced to a source

Primary Source

Publisher
NVIDIA
Published
2022-02-09

Data Quality

Spec confidence
Official
Clock basis
Boost
Core precisions with figures
4 of 9
Normalization
All values in TFLOPS

Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.

Browse GPU Cloud Providers

Similar GPUs

NVIDIA RTX A6000 48GB

PCIe · 2020
FP32: 39 TFLOPS
Compare vs RTX A6000

NVIDIA A40 48GB

PCIe · 2021
FP32: 37 TFLOPS
Compare vs A40

NVIDIA A10G 24GB

PCIe · 2022
FP32: 35 TFLOPS
Compare vs A10G

Frequently Asked Questions

How many TFLOPS does the NVIDIA A16 have?

The NVIDIA A16 delivers 4.5 TFLOPS FP32 and 18 TFLOPS FP16 at peak. Flopper does not currently have a verified FP8 throughput figure for it.

What is the power consumption of the NVIDIA A16?

Flopper does not currently have a verified TDP (Thermal Design Power) figure for the NVIDIA A16.

How much memory does the NVIDIA A16 have?

The NVIDIA A16 is equipped with 16 GB of memory with 200 GB/s of memory bandwidth.

What architecture is the NVIDIA A16 based on?

The NVIDIA A16 is based on the Ampere architecture, launched in 2021.

Stay Updated on GPU Releases

Get notified when new GPUs are added or specifications are updated.

Loading verification...

No spam, unsubscribe anytime.

Back to GPUs