AWS sells Inferentia2 only as EC2 Inf2 capacity, never as a component, so there is no board, no module standard and no list price, and form_factor is deliberately empty as it is on the Trainium rows. AWS has never published a TDP, a process node, a die size or a clock for any Neuron chip, so this row carries none of them and none should be added from third-party estimates. Two figures on this row do not behave the way a GPU ladder usually does and both are correct as published: FP8 equals FP16, BF16 and TF32 at 190 TFLOPS because AWS quotes one figure for the whole datatype group rather than doubling at FP8, and INT8 at 380 TOPS is the only rung above it. The 47.5 TFLOPS of FP32 is a Tensor Engine rate, not a vector rate; the NeuronCore-v2 Vector and Scalar engines are rated at 2.3 and 2.9 TFLOPS of FP32 each, which is an order of magnitude lower.
AWS logo

AWS Inferentia2

Type: ASICArchitecture: NeuronCore-v2Released: 2022Spec confidence: Official
FP8 (dense)
190
TFLOPS
FP32
48
TFLOPS
Memory
32 GB
HBM
Bandwidth
820 GB/s
memory

Overview

Inferentia2 is the inference half of the AWS Neuron line, the chip that sits under EC2 Inf2 while Trainium sits under Trn. Each chip is two NeuronCore-v2 cores over 32 GB of HBM at 820 GB/s, rated at 190 TFLOPS across FP16, BF16, cFP8 and TF32, 380 TOPS of INT8 and 47.5 TFLOPS of FP32. The generation is where AWS added FP32, TF32 and its configurable FP8 format to a line that had only done FP16, BF16 and INT8, and where it first put a chip-to-chip interconnect on an inference part: NeuronLink-v2 at 192 GB/s per chip, which is what lets a model larger than 32 GB be sharded across the twelve chips of an inf2.48xlarge instead of being confined to one. Against a contemporary GPU the compute is modest and the memory is small, and that is the point. Inferentia2 is sold as cost per token rather than peak throughput, and it is only available through EC2, never as a component. AWS publishes no power figure, no process node and no die size for any Neuron chip, so none is shown here.

Performance

Peak theoretical throughput by precision type

PrecisionPeak
FP64
No verified data available
FP32
32-bit floating point
48TFLOPS
TF32
TensorFloat-32
190TFLOPS
BF16
Brain Float 16
190TFLOPS
FP16
16-bit floating point
190TFLOPS
FP8
8-bit floating point
190TFLOPS
FP6
No verified data available
FP4
No verified data available
INT8
8-bit integer
380TOPS

Specifications

Architecture

NeuronCore-v2

Form Factor

No verified data available

Launch Year

2022

Process Node

No verified data available

Memory

32 GB HBM

Bandwidth

820 GB/s

TDP

No verified data available

Max power

No verified data available

Interconnect

192 GB/s NeuronLink-v2

direction and scope not stated by vendor

NeuronCore-v2

2

Spec Confidence

Official

Full Specifications

Compute Engine
NeuronCore-v2 2
Memory
Memory 32 GB
Memory Type HBM
Bandwidth 820 GB/s
Interface Width No verified data available
Interconnect & I/O
GPU-to-GPU NeuronLink-v2
Interconnect Bandwidth 192 GB/s direction and scope not stated by vendor
Power & Thermal
TDP No verified data available
Max power No verified data available
General
Form Factor No verified data available
Architecture NeuronCore-v2
Process Node No verified data available
Launch Year 2022

Datasheet & Resources

Flopper Datasheet

AWS Inferentia2 specifications, generated from our database. Printable.

Inferentia2 Architecture (AWS Neuron Documentation)

AWS · Latest version

View

Data Provenance

Every figure traced to a source

Primary Source

Publisher
AWS
Published
No verified data available

Data Quality

Spec confidence
Official
Clock basis
Boost
Core precisions with figures
6 of 9
Normalization
All values in TFLOPS

Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.

Browse GPU Cloud Providers

Similar GPUs

AWS Trainium

· 2020
FP32: 48 TFLOPS
Compare vs Trainium

Frequently Asked Questions

How many TFLOPS does the AWS Inferentia2 have?

The AWS Inferentia2 delivers 48 TFLOPS FP32, 190 TFLOPS FP16 and 190 TFLOPS FP8 at peak.

What is the power consumption of the AWS Inferentia2?

Flopper does not currently have a verified TDP (Thermal Design Power) figure for the AWS Inferentia2.

How much memory does the AWS Inferentia2 have?

The AWS Inferentia2 is equipped with 32 GB of memory with 820 GB/s of memory bandwidth.

What architecture is the AWS Inferentia2 based on?

The AWS Inferentia2 is based on the NeuronCore-v2 architecture, launched in 2022.

Stay Updated on GPU Releases

Get notified when new GPUs are added or specifications are updated.

Loading verification...

No spam, unsubscribe anytime.

Back to GPUs

© 2026 Flopper.io - Compare the GPUs Powering AI