AWS Inferentia2
Overview
Inferentia2 is the inference half of the AWS Neuron line, the chip that sits under EC2 Inf2 while Trainium sits under Trn. Each chip is two NeuronCore-v2 cores over 32 GB of HBM at 820 GB/s, rated at 190 TFLOPS across FP16, BF16, cFP8 and TF32, 380 TOPS of INT8 and 47.5 TFLOPS of FP32. The generation is where AWS added FP32, TF32 and its configurable FP8 format to a line that had only done FP16, BF16 and INT8, and where it first put a chip-to-chip interconnect on an inference part: NeuronLink-v2 at 192 GB/s per chip, which is what lets a model larger than 32 GB be sharded across the twelve chips of an inf2.48xlarge instead of being confined to one. Against a contemporary GPU the compute is modest and the memory is small, and that is the point. Inferentia2 is sold as cost per token rather than peak throughput, and it is only available through EC2, never as a component. AWS publishes no power figure, no process node and no die size for any Neuron chip, so none is shown here.
Performance
Peak theoretical throughput by precision type
| Precision | Peak |
|---|---|
FP64 | No verified data available |
FP32 32-bit floating point | 48TFLOPS |
TF32 TensorFloat-32 | 190TFLOPS |
BF16 Brain Float 16 | 190TFLOPS |
FP16 16-bit floating point | 190TFLOPS |
FP8 8-bit floating point | 190TFLOPS |
FP6 | No verified data available |
FP4 | No verified data available |
INT8 8-bit integer | 380TOPS |
Specifications
Architecture
NeuronCore-v2
Form Factor
No verified data available
Launch Year
2022
Process Node
No verified data available
Memory
32 GB HBM
Bandwidth
820 GB/s
TDP
No verified data available
Max power
No verified data available
Interconnect
192 GB/s NeuronLink-v2
direction and scope not stated by vendor
NeuronCore-v2
2
Spec Confidence
Official
Full Specifications
| Compute Engine | |
|---|---|
| NeuronCore-v2 | 2 |
| Memory | |
| Memory | 32 GB |
| Memory Type | HBM |
| Bandwidth | 820 GB/s |
| Interface Width | No verified data available |
| Interconnect & I/O | |
| GPU-to-GPU | NeuronLink-v2 |
| Interconnect Bandwidth | 192 GB/s direction and scope not stated by vendor |
| Power & Thermal | |
| TDP | No verified data available |
| Max power | No verified data available |
| General | |
| Form Factor | No verified data available |
| Architecture | NeuronCore-v2 |
| Process Node | No verified data available |
| Launch Year | 2022 |
Datasheet & Resources
Data Provenance
Every figure traced to a source
Primary Source
- Publisher
- AWS
- Published
- No verified data available
Data Quality
- Spec confidence
- Official
- Clock basis
- Boost
- Core precisions with figures
- 6 of 9
- Normalization
- All values in TFLOPS
Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.
Browse GPU Cloud ProvidersSimilar GPUs
Frequently Asked Questions
How many TFLOPS does the AWS Inferentia2 have?
The AWS Inferentia2 delivers 48 TFLOPS FP32, 190 TFLOPS FP16 and 190 TFLOPS FP8 at peak.
What is the power consumption of the AWS Inferentia2?
Flopper does not currently have a verified TDP (Thermal Design Power) figure for the AWS Inferentia2.
How much memory does the AWS Inferentia2 have?
The AWS Inferentia2 is equipped with 32 GB of memory with 820 GB/s of memory bandwidth.
What architecture is the AWS Inferentia2 based on?
The AWS Inferentia2 is based on the NeuronCore-v2 architecture, launched in 2022.
Stay Updated on GPU Releases
Get notified when new GPUs are added or specifications are updated.
No spam, unsubscribe anytime.