NVIDIA GPU Architectures Explained

Understanding the Evolution from Pascal to Blackwell

NVIDIA's GPU architectures represent major technology platforms that define entire generations of products. Each architecture introduces fundamental innovations in processing capabilities, memory systems, and specialized AI acceleration—spanning from consumer graphics cards to datacenter supercomputers.

Understanding the difference between an architecture (like Hopper), a product (like H100), and a system (like Grace Hopper Superchip) is essential for making informed decisions about GPU infrastructure. This guide breaks down NVIDIA's naming conventions, explains key concepts like Tensor Cores and NVLink, and traces the evolution from Kepler (2012) to today's cutting-edge Blackwell architecture.

What You'll Learn

  • How to decode NVIDIA's naming system (Architecture vs. Product vs. Superchip)
  • Key innovations in each architecture generation
  • Essential concepts: Tensor Cores, multi-die designs, NVLink, and Transformer Engine
  • Side-by-side technical comparisons across all datacenter architectures
Five generations of GPU modules, from bare board to liquid-cooled flagship
Five generations of GPU modules, from bare board to liquid-cooled flagship.
Architectures
41
Datacenter GPU generations
Latest
Blackwell
208B transistors, 2-die design
GPUs Tracked
243
Across all architectures

Decoding NVIDIA's Naming

Architecture vs Product vs Superchip — here's the difference in plain English

Architecture

The blueprint defining a GPU family's core technology and capabilities

Hopper • Blackwell • Ampere

Product

A specific GPU model built on an architecture — one arch, many products

H100 • H200 • B200

Superchip

CPU + GPU unified with ultra-fast interconnects for seamless data sharing

Grace Hopper • Grace Blackwell

GPU Architecture Evolution Timeline

Click on any architecture to learn more

Datacenter
Consumer
Superchip

Essential GPU Concepts

The building blocks that make modern AI acceleration possible

What is a Superchip?

A superchip combines different processor types (CPU + GPU) into one unified system.

NVIDIA Superchips like Grace Hopper integrate an ARM-based Grace CPU with a Hopper GPU using high-bandwidth NVLink-C2C interconnects. This creates a unified memory architecture where CPU and GPU can access the same memory space with 900GB/s bandwidth - 7x faster than PCIe Gen5. This eliminates memory copy overhead and enables seamless data sharing between processors.

Example: Grace Hopper (GH200) = Grace CPU + Hopper GPU via NVLink-C2C

What are Tensor Cores?

Specialized hardware units optimized for matrix operations in AI workloads.

Tensor Cores are specialized processing units designed to accelerate matrix multiplication operations fundamental to deep learning. Each generation adds new capabilities: Gen 1 (Volta) introduced FP16, Gen 2 (Turing) added INT8, Gen 3 (Ampere) brought TF32 and structural sparsity, Gen 4 (Hopper/Ada) added FP8 and Transformer Engine, and Gen 5 (Blackwell) introduced FP4 precision. They deliver 10-20x speedup over CUDA cores for AI training and inference.

Example: H100 has 4th-gen Tensor Cores supporting FP64, TF32, BF16, FP16, FP8, INT8

Multi-Die GPU Architecture

Using multiple silicon dies connected together to create one massive GPU.

Modern chip manufacturing has physical size limits (reticle limit). Multi-die architecture overcomes this by connecting 2+ GPU dies via ultra-fast interconnects, making them appear as a single GPU to software. Blackwell pioneered this with its 2-die design reaching 208B transistors - impossible in a single die. The interconnect bandwidth must be extremely high to avoid bottlenecks between dies.

Example: Blackwell B200 uses 2 dies connected at 10TB/s, totaling 208B transistors

NVLink: GPU-to-GPU Interconnect

High-speed interconnect for direct GPU-to-GPU communication.

NVLink is NVIDIA's proprietary high-bandwidth, low-latency interconnect for multi-GPU systems. It allows GPUs to directly share memory and communicate without going through the CPU or PCIe bus. NVLink bandwidth has evolved: Gen 1 (Pascal) at 160GB/s, Gen 2 (Volta) at 300GB/s, Gen 3 (Ampere) at 600GB/s, Gen 4 (Hopper) at 900GB/s, and Gen 5 (Blackwell) at 1800GB/s. Essential for training large models across multiple GPUs.

Example: H100 with 900GB/s NVLink enables efficient 8-GPU training clusters

Explore NVIDIA Architectures

Click any architecture to dive deep into specs, peak performance, and GPU models

CDNA 6

2027

datacenter
HBM4E memory on a 2nm process, powering the AMD Helios 500 rack-scale platform

CDNA 6 is the sixth generation of AMD's data center compute architecture, named by AMD at CES 2026 and scheduled to arrive with the Instinct MI500 Series in 2027. AMD has stated three things about it and no more: it is built on an advanced 2nm process, it uses HBM4E memory, and it succeeds the CDNA 5 architecture that debuted with the Instinct MI400 series in 2026. AMD has published no compute unit count, no clock speed, no transistor count, no memory capacity or bandwidth, no board power and no throughput figure at any precision for any CDNA 6 part. Its single performance claim is a projection that an MI500-powered AI rack will deliver up to a thousand times the AI performance of an MI300X platform, a comparison AMD makes between a rack and a platform without naming a precision on either side, and which therefore cannot be reduced to a per-accelerator figure. CDNA 6 silicon will power the AMD Helios 500 rack-scale system alongside EPYC "Verano" CPUs and Pensando "Como" and "Monza" networking, and AMD has already named its successor, the MI600 Series in the Helios 600 rack.

Process
2nm
Tensor Cores
Gen 6

CDNA 5

2026

datacenter
Wave32 Work Group Processors, 3D hybrid bonded compute dies, HBM4 and UALink over Ethernet scale-up

Fifth generation AMD CDNA architecture, introduced with the Instinct MI400 series in 2026. CDNA 5 moves the compute front end to Work Group Processors with Wave32 execution, stacks 3D hybrid bonded compute dies on a CoWoS-L package, and splits the design into eight TSMC N2 Accelerator Complex Dies plus two I/O dies and two Fabric and Cache dies on N3. It is the first Instinct generation on HBM4 and the first to use UALink over Ethernet for rack scale-up, and it quadruples dense OCP MXFP4 and MXFP8 matrix throughput per GPU versus CDNA 4.

Transistors
320B
Max HBM
432 GB
Process
TSMC N2 (XCD) / N3 (IOD, FCD)
Tensor Cores
Gen 5
Design
8-Die
Huawei logo

Da Vinci v3

2026

datacenter
Huawei's first generation with native low-precision compute, adding FP8 and MXFP4 alongside its own HiF8 format, and the first to use Huawei-designed memory, HiBL 1.0 for prefill and HiZQ 2.0 for decode

Da Vinci v3 is the Ascend 950 generation and the first Huawei architecture with native FP8 and 4-bit compute. It is also the first split into two memory-specialised parts: the Ascend 950PR pairs 128 GB of Huawei's own HiBL 1.0 at 1,600 GB/s for prefill and recommendation work, while the Ascend 950DT pairs 144 GB of HiZQ 2.0 at 4,000 GB/s for decode. Both carry 2,000 GB/s of Unified Bus 2.0 scale-up. The Atlas 350 accelerator card is built on the 950PR chip as a de-rated bin, at seven eighths of its memory and bandwidth. Huawei has published no process node, transistor count or die size for this generation, and the parts are export-controlled to the Chinese domestic market.

Max HBM
144 GB
Max TDP
900W
Qualcomm logo

Dragonfly

2026

datacenter
768 GB of LPDDR per accelerator card, the highest per-accelerator memory capacity of any announced AI accelerator, trading memory bandwidth for capacity and cost in rack-scale inference

Dragonfly is Qualcomm's brand for its data center AI silicon, introduced with the AI200 and AI250 rack-scale inference accelerators announced on 27 October 2025 and extended by a third-generation AI300 on Qualcomm's published annual cadence. Its defining decision is memory: Dragonfly accelerators carry LPDDR rather than HBM, reaching 768 GB per card where contemporary HBM parts reach 144 GB to 432 GB, and accepting far lower bandwidth in exchange for capacity, cost and supply availability. The design target is large-model inference economics rather than peak throughput, and Qualcomm markets it on total cost of ownership and performance per dollar per watt rather than on FLOPS. Qualcomm builds on its NPU heritage from the Cloud AI 100 family and Snapdragon, but has published no microarchitecture name, no NPU generation number, no AI-core count, no clock speed, no process node and no throughput figure of any kind for any Dragonfly part. The second-generation AI250 adds a near-memory compute architecture that Qualcomm claims delivers more than ten times the effective memory bandwidth of the AI200.

Maia 200

2026

datacenter
Inference-first tile and cluster hierarchy with narrow-precision FP4 and FP8 matrix datapaths, 272MB of software-managed on-die SRAM, and an on-die NIC driving a switchless two-tier Ethernet scale-up fabric

Maia 200 is Microsoft's second-generation in-house AI accelerator architecture and its first silicon and system platform built specifically for inference rather than training. Compute is organised as a hierarchy: a tile pairs a Tile Tensor Unit for matrix multiply with a Tile Vector Processor for programmable SIMD work, backed by multi-banked Tile SRAM, a tile DMA engine and a lightweight Tile Control Processor; tiles compose into clusters that share a large Cluster SRAM and a cluster DMA subsystem staging traffic to co-packaged HBM; clusters compose into the SoC. Both SRAM tiers are fully software managed, so compilers and kernels pin working sets on die rather than relying on cache heuristics, and tile and SRAM redundancy schemes are built in for yield. The Tile Tensor Unit is optimised for FP8, FP6 and FP4 matrix multiplication including mixed FP8 activations against FP4 weights, with hardware casting between storage and compute types at line rate. Scale-up is Ethernet rather than a proprietary fabric: an on-die NIC provides 1.4 TB/s unidirectional bandwidth, Microsoft's AI Transport Layer protocol adds packet spraying, multipath routing and congestion-resistant flow control, groups of four accelerators form a switchless Fully Connected Quad, and a switched second tier extends the domain to 6,144 accelerators.

Transistors
140B
Max HBM
216 GB
Process
TSMC N3
Max TDP
750W

Rubin

2026

datacenter
NVFP4 microscaled 4-bit tensor format plus Tensor Core emulated SGEMM and DGEMM, which lift FP32 to 400 TFLOPS and FP64 to 200 TFLOPS on a GPU whose native ALU rates are 130 and 33 TFLOPS

Rubin is NVIDIA's datacenter GPU architecture for 2026, succeeding Blackwell and Blackwell Ultra and named after the astronomer Vera Rubin. A Rubin GPU is a 336 billion transistor part built from two reticle-limited compute dies joined by the NVIDIA High-Bandwidth Interface, carrying 224 streaming multiprocessors and 896 Tensor Cores, paired with 288 GB of HBM4 at 22 TB/s and connected by sixth-generation NVLink at 3.6 TB/s per GPU. It pairs with the 88-core Vera CPU over NVLink-C2C to form the Vera Rubin Superchip, and 72 GPUs with 36 CPUs form the Vera Rubin NVL72 rack. Rubin introduces NVFP4, a microscaled four-bit format with block-level scaling, and a third-generation Transformer Engine with hardware-accelerated adaptive compression. It is also the first NVIDIA architecture to offer both native and Tensor Core emulated FP32 and FP64, so an SGEMM or DGEMM workload sees roughly three and six times the native ALU throughput respectively. The generation also included Rubin CPX, a monolithic GDDR7 part for the context phase of long-context inference, which NVIDIA removed from its roadmap at GTC 2026, and is followed in 2027 by Rubin Ultra.

Transistors
336B
Max HBM
288 GB
Design
2-Die

Xe3P

2026

datacenter
LPDDR5X instead of HBM for high-capacity air-cooled inference, with native FP4 and MXFP4 microscaling formats through to FP64

Xe3P is the performance-optimised variant of Intel's Xe3 graphics architecture. Its first announced datacenter application is Crescent Island, an inference accelerator that pairs Xe3P with LPDDR5X rather than HBM in order to reach very large memory capacity in a 350 W air-cooled PCIe card. Intel has described the datatype range as spanning native FP4 and MXFP4 microscaling formats through to FP64, but has published no throughput, clock, core count or process node figures for any Xe3P part.

Max TDP
350W
T-Head logo

Zhenwu

2026

datacenter
ICN, a self-developed inter-chip network that gives every accelerator its own independent ports and memory-semantic addressing into its neighbours, scaled through an ICN Switch to full-bandwidth domains of 64, 128 and 1024 cards

Zhenwu is T-Head's datacenter AI accelerator line, built around a self-developed parallel computing architecture and the ICN inter-chip network, and paired with T-Head SAIL, the company's own software stack. Alibaba describes every member as a 训推一体 part, one chip for both training and inference, and the line is sold as Alibaba Cloud capacity rather than as a component. Three generations exist. The 810E carries 96 GB of HBM2e at 2.7 TB/s with seven ICN ports and 700 GB/s of inter-chip bandwidth. The M890 raises that to 144 GB, eight ports and 800 GB/s, adds FP8 and FP4 to a range that already ran from FP32 through BF16 and FP16, and introduces the ICN Switch, which makes 64 cards look like one flat full-bandwidth domain. The V900 carries 216 GB and 1200 GB/s, deepens the tensor units' FP8 and FP4 instruction precision, adds configurable scaling-factor formats and block sizes for MXFP8 and MXFP4, and scales the same Scale-Up architecture to 1024 chips through two layers of ICN Switch. The defining absence is arithmetic: T-Head has never published a throughput figure for any Zhenwu part, quoting only generational ratios of three times the previous chip.

Apple M5

2025

consumer
Process
3nm

Blackwell Ultra

2025

datacenter
Enhanced Blackwell with 288GB HBM3e

Enhanced variant of Blackwell architecture featuring increased HBM capacity (288GB) and improved FP4 performance for the most demanding AI workloads.

Transistors
208B
Max HBM
288 GB
Process
TSMC 4NP
Max TDP
1.4KW
Tensor Cores
Gen 5
Design
2-Die
EF

GCU-CARA 4

2025

datacenter
A fourth-generation compute unit built on Enflame's own instruction set rather than on a GPGPU model, whose distinguishing feature is native FP8 arithmetic rather than FP8 emulated on wider hardware

GCU-CARA is Enflame's name for the accelerated compute unit at the centre of its cloud AI chips, and the fourth generation is the one behind the 邃思 400 chip and the L600 module. Enflame's IPO prospectus describes the line as a deliberate refusal of the prevailing design: "公司未跟随英伟达的 GPGPU 架构,而是基于自主指令集", the company did not follow NVIDIA's GPGPU architecture but built on its own instruction set, with GCU-CARA answering the Tensor Core and a separate technology, GCU-LARE, answering NVLink. The instruction set covers compute, control, cache access and data synchronisation, and the microarchitecture has now iterated four times. What the fourth generation adds is native FP8: Enflame describes itself as one of the few Chinese vendors whose silicon supports the datatype in hardware rather than by conversion, which is the same distinction that separates the MetaX C700's FP8 compute from the C600's FP8 conversion instructions. Alongside it sits the fourth-generation GCU-LARE interconnect, which Enflame says supports direct chip-to-chip topologies past the bandwidth and latency limits of PCIe and scales to superpods and clusters of ten thousand cards and beyond. The generation ladder Enflame publishes runs 邃思 1.0 in 2019, 邃思 2.0 and 2.5 in 2021, 邃思 320 in 2024 and 邃思 400 in 2025. Enflame's own filing spells the unit both GCU-CARA, in its product table, and GCU-CARE, in its core-technology section; the table spelling is used here.

RDNA 4

2025

consumer
Second-generation AI accelerators with native FP8 in both E5M2 and E4M3 encodings and 2:1 structured sparsity, taking the matrix rate from twice the vector rate on RDNA 3 to four times it, and making AMD publish a dense and a sparse figure side by side for every matrix precision

RDNA 4 is AMD's 2025 graphics architecture and the first Radeon generation with AI hardware worth quoting next to a datacenter part. Its second-generation AI accelerators add native FP8 in both E5M2 and E4M3 encodings and 2:1 structured sparsity, neither of which RDNA 3 had. On the 64 compute unit configuration that means 195 TFLOPS of dense FP16 matrix and 389 TFLOPS of dense FP8 on the Radeon RX 9070 XT, four and eight times its 48.7 TFLOPS FP32 vector rate respectively, where RDNA 3 stopped at twice. AMD prints the dense and the structured-sparsity figure side by side on every product page in the family, which makes it one of the easier vendor tables to read correctly. Memory stays GDDR6 rather than HBM, so capacity tops out at 32 GB on the Radeon AI PRO R9700 and bandwidth at 640 GB/s on a 256-bit bus, and that is the binding limit on the family for local model work rather than compute. AMD publishes no FP64 and no BF16 rate for any RDNA 4 part, and no lithography node for any of them either. There is no datacenter RDNA 4 product; Instinct continues on the separate CDNA line.

Transistors
53.9B
Max TDP
304W
Tensor Cores
Gen 2
MetaX logo

XCORE 1.5

2025

datacenter
The second iteration of MetaX's compute GPU IP, moving to HBM3e and an OAM module at kilowatt power to compete for frontier training work in the Chinese domestic market

XCORE 1.5 is the second generation of MetaX's compute GPU IP, named as such in the company's STAR Market filing, and this database holds it as the C600. The part moves the line to HBM3e and an OAM module: 144 GB and up to 1,000 W, with MetaXLink for scale-up. That capacity is a press-corroborated figure sitting inside MetaX's own published floor of more than 96 GB, and it is flagged as such on the GPU row rather than presented as a datasheet value. Memory bandwidth, process node and throughput figures are all unpublished, and a widely circulated set of C600 numbers is disqualified because its power figure contradicts MetaX's own.

Max HBM
144 GB
Max TDP
1.0KW
MetaX logo

XCORE 2.5

2025

datacenter
The next compute instruction set on MetaX's own published ladder, funded and in design but not taped out, with no specification of any kind released

XCORE 2.5 is the third generation of MetaX's compute GPU instruction set, named in the company's STAR Market filing as the successor built on the XCORE 1.5 of the C600. This database holds it as the C700, and every numeric column on both rows is empty because MetaX has published nothing: no throughput at any precision, no memory capacity, type or bandwidth, no power figure, no process node and no interconnect. At the date of the filing the project was in physical implementation and the chip had not been manufactured. The widely repeated claim that the C700 approaches an H100 is a qualitative positioning sentence in an IPO document with no number attached, and nothing here derives from it. Two naming traps: XCORE 2.0 is MetaX's RENDER instruction set and is not the compute successor to XCORE 1.5, and MetaX never writes C700 and XCORE 2.5 in the same sentence, so the pairing is an identification from the company's own ISA ladder rather than a MetaX label. The year on both rows is the disclosure year, not a launch year.

Apple M4

2024

consumer

Fourth-generation Apple Silicon SoC on second-gen 3nm process. Enhanced GPU performance and power efficiency with LPDDR5X support.

Process
3nm

Battlemage

2024

both
XMX matrix engines widened to 2,048 bits and rebuilt around a native SIMD16 vector engine, putting workstation-class INT8 inference throughput inside a 190 W consumer card and a 70 W half-height one

Battlemage is Intel's second generation of discrete Arc graphics, announced on 3 December 2024 and built on the Xe2 microarchitecture at TSMC N5. Its Xe-core is rebuilt around native SIMD16 execution: eight 512-bit vector engines and eight 2,048-bit XMX matrix engines per core, with 256 KB of shared L1 and scratchpad memory, three-way co-issue of floating point, integer and matrix work, and XMX datatype support spanning INT2, INT4, INT8, FP16, BF16 and TF32. Intel claims 70 per cent more performance per Xe-core and 50 per cent more performance per watt than Alchemist. The desktop cards use the BMG-G21 die: five render slices, 20 Xe-cores, 160 XMX engines, 20 ray tracing units, 18 MB of L2 cache and a 192-bit GDDR6 interface. The same die underpins the Arc Pro B50 and B60 workstation cards, which is why this architecture is marked for both consumer and datacenter use; the Arc Pro B60 is sold specifically for local large language model inference and is the building block of Intel's Project Battlematrix workstations, which put eight of them and 192 GB of memory in one Xeon chassis. Intel publishes one throughput figure per Battlemage card, peak INT8 TOPS on XMX with dense models, and no FP32, FP16, BF16, TF32 or FP64 rate for the consumer parts. Note that Xe2 is the microarchitecture and Battlemage the desktop die family, the same relationship Alchemist has to Xe-HPG in this database; no separate Xe2 row exists yet.

Process
5nm
Max TDP
290W

Blackwell

2024

both
First 2-die GPU design with FP4 precision

Next-generation 2-die GPU architecture with 208B transistors, 5th-gen Tensor Cores supporting FP4 precision, and breakthrough AI inference performance.

Transistors
208B
Max HBM
192 GB
Process
TSMC 4NP
Max TDP
1.2KW
Tensor Cores
Gen 5
Design
2-Die

RDNA 3.5

2024

consumer
A power-optimised revision of RDNA 3 built for integrated graphics, which on the Strix Halo parts is paired with a 256-bit LPDDR5X memory controller so an integrated GPU finally gets 256 GB/s of bandwidth and up to 96 GB of the system memory addressable as video memory

RDNA 3.5 is AMD's integrated-graphics revision of RDNA 3, announced at Computex on 2 June 2024 alongside the Ryzen AI 300 series and the Radeon 800M graphics that ship inside it. It is not a discrete generation: there is no RDNA 3.5 graphics card, and AMD publishes no compute-unit-level architectural detail for it in the way it does for RDNA 3 and RDNA 4. What makes it matter for this catalog is the Strix Halo configuration, where AMD pairs up to 40 RDNA 3.5 compute units with a 256-bit LPDDR5X memory controller. That gives an integrated GPU 256 GB/s of memory bandwidth and, through AMD's Variable Graphics Memory, up to 96 GB of a 128 GB unified pool addressable as video memory, which is a great deal more capacity than any consumer discrete card offers and the reason these parts get bought for running large models locally. AMD publishes no FP32, FP16, BF16 or INT8 throughput figure for any RDNA 3.5 integrated GPU, and no AI Accelerator or Ray Accelerator count either, so the only compute number available for these parts is a vector rate computed on AMD's own arithmetic. The TOPS figures AMD does publish for Ryzen AI describe the separate XDNA 2 neural processor or the whole SoC, not the graphics.

Process
4nm
Max TDP
120W
Tenstorrent logo

Wormhole

2024

both
Ethernet moved onto the die, so a Tensix mesh can be extended across cards and across chassis without a switch or a proprietary fabric, and a two-chip card is programmed as one continuous grid

Wormhole is the second generation of Tenstorrent's Tensix architecture, following Grayskull and preceding Blackhole. A Tensix core is not a shader: it is a matrix engine and a SIMD vector unit wrapped around five small RISC-V cores that handle unpack, math, pack and two network-on-chip interfaces, with roughly 1.5 MB of local SRAM each. The defining change in Wormhole is that Ethernet is on the die, so the mesh of Tensix cores extends off-chip over standard Ethernet rather than a proprietary fabric, which is how a two-processor n300 card presents as one continuous grid and how cards are meshed to each other through their QSFP-DD ports without going through the host. Tenstorrent publishes three compute figures for Wormhole, FP8, FP16 and BFP8, and marks none of them sparse, because the hardware has no structured sparsity. It also lists a long tail of supported formats it publishes no rate for, including BF16, TF32, INT8, INT32, several block-float widths and its own VTF19. The whole software stack, tt-metal and tt-forge, is open source, which together with a card price near a thousand dollars is most of the reason these parts get bought.

Max TDP
300W

Apple M3

2023

consumer

Third-generation Apple Silicon SoC built on TSMC 3nm process. Features Dynamic Caching, hardware-accelerated ray tracing, and mesh shading.

Process
3nm
Qualcomm logo

Cloud AI 100

2023

datacenter
576 MB of on-die SRAM across 64 software-managed AI cores, so that inference weights and activations stay on the die rather than crossing to DRAM

Cloud AI 100 is Qualcomm's first data center AI architecture, a fixed-function inference design built on a 7 nm process and sold as PCIe cards rather than as a rack-scale system. Its defining choice is on-die memory: each AI core owns 9 MB of software-managed SRAM, giving 576 MB on the largest card, which is more local scratchpad memory than any other non-wafer-scale accelerator carries and which lets inference workloads avoid DRAM traffic entirely for many layers. Capacity comes from LPDDR4x rather than HBM, trading bandwidth for cost and power in a 150 W envelope. Qualcomm publishes throughput for exactly two datatypes, INT8 and FP16, and no architectural detail beyond the AI core count and its SRAM allocation: there is no published clock speed, no MAC width, no die size and no transistor count for any part in the family. The line spans five SKUs cut from the same core, differing in how many cores are enabled and at what clock, from a 16-core 75 W card to a 64-core 150 W card. Qualcomm has never applied its later Dragonfly data center brand to this family, and never uses the word Hexagon in any Cloud AI 100 document. It remains the only Qualcomm data center accelerator that has shipped and can be rented, its Dragonfly successors having been announced with no published throughput figures and no availability.

Process
7nm
Meta logo

MTIA

2023

datacenter
High-velocity modular chiplet design: a new accelerator generation roughly every six months, built from reusable compute, network and SoC chiplets, with an inference-first focus and a PyTorch-native software stack

MTIA, the Meta Training and Inference Accelerator, is Meta's family of in-house AI accelerators, developed in close partnership with Broadcom and deployed only inside Meta's own data centers. The architecture is organised around a grid of processing elements. Each PE contains two RISC-V vector cores, a Dot Product Engine for matrix multiplication, a Special Function Unit for activations and elementwise operations, a Reduction Engine for accumulation and inter-PE communication, and a DMA engine feeding local scratch memory. Chips are assembled from modular chiplets, which is what lets Meta ship a new generation on a roughly six-month cadence rather than a two-year one: MTIA 300 uses one compute chiplet with two network chiplets, MTIA 400 combines two compute chiplets to double compute density, and MTIA 500 moves to a two-by-two arrangement of smaller compute chiplets surrounded by HBM stacks, two network chiplets and an SoC chiplet providing PCIe connectivity to the host. The design philosophy is explicitly inference-first and cost-driven rather than peak-FLOPS driven, with built-in NIC chiplets, dedicated message engines that offload communication collectives, and near-memory compute for reduction-based collectives. Later generations add hardware acceleration for attention and mixture-of-experts bottlenecks and custom low-precision data types intended to raise throughput without hurting model quality. Across the generation Meta states that HBM bandwidth rises 4.5 times and compute rises 25 times, though that 25x figure compares MTIA 300 at 8-bit precision with MTIA 500 at 4-bit precision rather than measuring like for like. Software is PyTorch native, with a graph compiler, a Triton compiler and eager-mode runtime support, which Meta treats as central to getting internal workloads onto the hardware quickly.

Max HBM
512 GB
Max TDP
1.7KW

Ada Lovelace

2022

both
Fourth-generation Tensor Cores with native FP8, paired with a large L2 cache and Shader Execution Reordering, giving datacenter inference cards Hopper-class low-precision throughput without HBM

Ada Lovelace is NVIDIA's 2022 graphics architecture and the generation that made FP8 inference cheap on GDDR6 cards. Its fourth-generation Tensor Cores add native FP8 support, and this database holds Ada parts across three market segments: the datacenter L-series (L40S, L40, L20, L4, L2), the RTX workstation cards (RTX 6000 Ada, RTX 5000 Ada, RTX 2000 Ada) and GeForce (RTX 4090, 4080, 4080 Super, 4070 Ti). The L40S is the family's inference workhorse at 48 GB of GDDR6, 864 GB/s, 91.6 TFLOPS FP32 and 733 TFLOPS dense FP8 in a 350 W PCIe card. The trade-off against a datacenter architecture is FP64: the L40S manages 1.45 TFLOPS against its own 91.6 TFLOPS FP32, roughly a sixty-fourth, so Ada is an inference and visualisation architecture rather than an HPC one. It also has no NVLink on any part in this database, and no HBM anywhere in the family.

Max TDP
450W
Tensor Cores
Gen 4

Alchemist

2022

both
Intel's first generation of discrete gaming GPUs in two decades, the ACM-G10 and ACM-G11 dies implementing Xe-HPG

Alchemist is the codename for Intel's first-generation Arc discrete GPU family, built on the Xe-HPG microarchitecture across two dies, ACM-G10 and ACM-G11. This database holds the Arc A770 with 16 GB of GDDR6 at 560 GB/s and a 225 W board power. Alchemist is recorded as a child of Xe-HPG: Xe-HPG is the microarchitecture and Alchemist is the die family that implements it, the same relationship the Data Center GPU Flex Series parts have to Xe-HPG.

Process
6nm
Max TDP
225W
Tensor Cores
Gen 1

Apple M2

2022

consumer

Second-generation Apple Silicon SoC. Up to 10-core GPU with improved performance per watt and higher memory bandwidth.

Process
5nm
Biren logo

Biren SPC

2022

datacenter
A chiplet GPGPU presented at Hot Chips 2022 with 2,300 GB/s of BLink scale-up, the most aggressive interconnect any Chinese accelerator had published at the time

Biren SPC is Biren Technology's first-generation general-purpose GPU architecture. The name is a label used by this database rather than a Biren one: the company calls it simply its first-generation GPGPU architecture and has never given it a name. The specified part is the BR100, presented at Hot Chips 2022 and for a while the most ambitious Chinese datacenter GPU design published anywhere: a 64 GB HBM2e OAM module at 1,600 GB/s and 550 W, rated at 256 TFLOPS FP32, 512 TFLOPS TF32, 1,024 TFLOPS BF16 and 2,048 TOPS INT8, with 2,300 GB/s of BLink scale-up bandwidth. Those figures come from Biren's own conference presentation rather than a datasheet, so they are recorded as a vendor claim. The inference-oriented BR110 shares this architecture and is also held here, but with no specifications at all: Biren has never published a single throughput, memory, power or process figure for it, and nothing on that row is inherited from the BR100. Read the generation in its commercial context, which is the largest fact about it: Biren has been under US export restrictions since October 2023, which cut it off from leading-edge foundry capacity, so shipped volume is very much smaller than the specification implies.

Max HBM
64 GB
Process
TSMC 7nm
Max TDP
550W

Hopper

2022

datacenter
Transformer Engine with FP8 precision

Fourth-generation datacenter GPU architecture featuring Transformer Engine, 4th-gen Tensor Cores with FP8 support, and enhanced NVLink for AI training and HPC workloads.

Transistors
80B
Max HBM
141 GB
Process
TSMC 4N
Max TDP
700W
Tensor Cores
Gen 4
AWS logo

NeuronCore-v2

2022

datacenter
Configurable FP8 (cFP8) and TF32 datatypes, a programmable GPSIMD engine for custom operators, and hardware support for dynamic shapes and control flow

NeuronCore-v2 is the second generation of the AWS Neuron compute core. Each core is a fully independent heterogeneous compute unit with four engines: a systolic-array Tensor Engine delivering over 90 TFLOPS of FP16 or BF16, a Vector Engine at 2.3 TFLOPS of FP32, a Scalar Engine at 2.9 TFLOPS of FP32, and a new GPSIMD Engine of eight fully programmable 512-bit vector processors that run general-purpose C code against the embedded on-chip SRAM, which is how custom operators are implemented. The generation introduces the configurable FP8 datatype AWS calls cFP8 alongside TF32, adds control flow, dynamic shapes and programmable rounding including stochastic rounding, and scales out over NeuronLink-v2. Two NeuronCore-v2 cores make up one Inferentia2 chip.

Max HBM
32 GB
Tensor Cores
Gen 2

RDNA 3

2022

consumer
The first chiplet consumer GPU, separating the graphics compute die from memory-cache dies, with dedicated AI accelerators added to the shader array

RDNA 3 is the graphics half of AMD's split roadmap, the consumer counterpart to CDNA 3, and the first consumer GPU architecture built from chiplets: a graphics compute die packaged with separate memory-cache dies. This database holds two RDNA 3 parts, the Radeon RX 7900 XTX (24 GB, 960 GB/s, 355 W) and the RX 7900 XT (20 GB, 800 GB/s, 315 W). They are included for cross-architecture comparison rather than as datacenter options, and the numbers show why: the 7900 XTX delivers 61.4 TFLOPS FP32 and 122.8 TFLOPS FP16 through its AI accelerators, a 2x step where a CDNA or Tensor Core part manages 8x or more, and its FP64 rate is 1.54 TFLOPS, about a fortieth of FP32. There is no HBM, no Infinity Fabric scale-up and no ECC HBM stack anywhere in the family.

Process
5nm
Max TDP
355W

Ampere

2020

both
Third-generation Tensor Cores introducing TF32 and structured 2:4 sparsity, plus Multi-Instance GPU, which partitions one physical accelerator into up to seven isolated instances

Ampere is NVIDIA's 2020 architecture and the widest family in this database, spanning the datacenter GA100 die and the consumer and professional GA10x dies. Its third-generation Tensor Cores introduced TF32, a 19-bit format that runs unmodified FP32 training code on the tensor path, and structured 2:4 sparsity, which doubles the quoted rate when a model is pruned to that pattern -- every value stored here is the DENSE figure. The A100 is the reference part: 80 GB of HBM2e at 2,039 GB/s in the SXM4 form factor, 19.5 TFLOPS FP32, 156 TFLOPS TF32, 312 TFLOPS FP16 and BF16, 624 TOPS INT8, and 9.7 TFLOPS FP64 at half the FP32 rate, with 600 GB/s of NVLink. Ampere also brought Multi-Instance GPU, and it is the generation the export-controlled A800 belongs to. Two nodes are used across the family, 7 nm for GA100 and 8 nm for GA10x, so no single process node is true of the architecture.

Transistors
54.2B
Max HBM
80 GB
Max TDP
450W
Tensor Cores
Gen 3

Apple M1

2020

consumer

First-generation Apple Silicon SoC. 8-core GPU based on Apple's custom GPU architecture with unified memory.

Process
5nm
Graphcore logo

Colossus MK2

2020

datacenter
An accelerator with no DRAM at all, putting 900 MB of compiler-managed SRAM inside the processor beside 1,472 independent MIMD cores, so that every core runs its own program against local memory instead of sharing a scheduler and a cache hierarchy

Colossus MK2 is the second generation of Graphcore's Intelligence Processing Unit, announced in July 2020 as the GC200. Each chip is 59.4 billion transistors on an 823 square millimetre TSMC 7nm die, organised as 1,472 IPU-Cores running 8,832 worker threads, and its defining property is that all of its working memory, 900 MB of it, is SRAM inside the processor. There is no HBM, no GDDR and no cache hierarchy: memory is addressed directly and scheduled by the Poplar compiler, and the tile-local arrangement is what gives the architecture its very high on-chip bandwidth. Cores are fully independent, which suits sparse and irregular graph work that a lockstep SIMT machine handles badly, and the trade is that a model larger than 900 MB has to be split across chips over the IPU-Link fabric or streamed from DDR4 on the host machine. Numerics follow Graphcore's AI-Float convention, which names the multiply and accumulate widths separately as FP16.16 and FP16.32, with stochastic rounding in hardware so that FP16 master weights hold their accuracy. Three parts share the generation. The GC200 is the original, rated at 250 TFLOPS of FP16.16 and 62.5 TFLOPS of FP32. The Bow IPU of March 2022 is the same architecture built as a wafer-on-wafer 3D stack, with a power delivery die bonded underneath that lets the logic clock at 1.85 GHz for 350 TFLOPS. The C600 of November 2022 is a revision with FP8 added, the only member sold as a PCIe card, reaching 560 TFLOPS of FP8 and 280 of FP16 at 185 W.

Transistors
59.4B
Process
TSMC 7nm

RDNA 2

2020

consumer
AMD Infinity Cache, a 128 MB on-die last-level cache that let a 256-bit GDDR6 bus feed a 4K-class GPU, alongside the first hardware ray accelerators in a Radeon part

RDNA 2 is AMD's 2020 graphics architecture, the generation that powers the PlayStation 5 and Xbox Series consoles as well as the Radeon RX 6000 and Radeon PRO W6000 desktop cards. Its defining idea is Infinity Cache: up to 128 MB of on-die last-level cache that absorbs enough traffic to let a 256-bit GDDR6 bus feed an 80 compute unit GPU, which is how the Radeon RX 6900 XT reaches 23 TFLOPS of FP32 on 512 GB/s of external bandwidth. It also brought the first hardware ray accelerators to Radeon, one per compute unit. For machine learning the important characteristic is a negative one: RDNA 2 has no matrix cores at all. Every performance figure AMD publishes for it is a vector figure, with FP16 running at exactly twice FP32 through packed math, and there is no BF16, FP8, INT8 or INT4 throughput published anywhere in the generation. Matrix hardware arrived on the Radeon side one generation later, with RDNA 3. AMD built RDNA 2 on TSMC 7 nm across three dies, Navi 21 at 26.8 billion transistors, Navi 22 at 17.2 billion and Navi 23 at 11.1 billion, and every part in the family uses GDDR6 rather than HBM.

Process
7nm
Max TDP
335W
Groq logo

Tensor Streaming Processor

2019

datacenter
A processor with no caches, no branch prediction, no reorder buffer and no external memory, where the compiler schedules every instruction issue cycle by cycle, so that execution time is known before the program runs and does not vary between runs

The Tensor Streaming Processor is the architecture Groq announced in November 2019 and later rebranded the LPU. It inverts the conventional layout: instead of a two-dimensional mesh of identical cores each holding a slice of every function, the TSP slices the functions themselves into parallel lanes that run the width of the die, so that data streams sideways through a memory slice, a vector slice, a matrix slice and a switch slice in turn. Nothing is scheduled by hardware. There are no caches to miss, no branch prediction, no reorder buffer and no external memory of any kind, only 230 MB of globally shared SRAM addressed directly at up to 80 TB/s, with the compiler placing every instruction on a known cycle. The result is determinism: the same program takes exactly the same number of cycles every time it runs, which is why Groq sells the design on tail latency rather than on peak throughput. The matrix unit is an integer array, and floating point arrives through Groq's TruePoint technique, which decomposes an FP16 or FP32 multiply into integer work on the same hardware; that is why the first-generation part is rated four times faster at INT8 than at FP16 rather than twice. The trade is capacity. A model larger than 230 MB has to be split across many chips, so the architecture spends its area on very high radix chip-to-chip links instead of on local memory. First silicon was a 25 by 29 millimetre 14nm die of 26.8 billion transistors running at a nominal 900 MHz.

Transistors
26.8B
Process
14nm

Turing

2019

both
Integer Tensor Core paths, INT8 and INT4, which turned quantised inference into the mainstream deployment format and made a 70 W single-slot accelerator viable

Turing is NVIDIA's 2018-2019 architecture, and this database holds it across the datacenter T4, the Quadro RTX workstation pair and the GeForce and Titan cards that introduced ray tracing. Its contribution to inference was integer precision: adding INT8 and INT4 paths to the Tensor Core took the T4 to 130 TOPS INT8 and 260 TOPS INT4 against 65 TFLOPS FP16 and 8.1 TFLOPS FP32, and it did so inside a 70 W single-slot passively cooled PCIe card with 16 GB of GDDR6 at 300 GB/s. That power envelope is the point: the T4 became the default inference card in dense server deployments for years. There is no NVLink on the T4 and no HBM anywhere in the family, and the 32 GB/s recorded on its row is its PCIe link, not a scale-up fabric. Turing is also the generation that introduced hardware ray tracing on the consumer side, and the GeForce RTX 20 series and TITAN RTX rows here carry that consumer side of the family.

Process
12nm
Max TDP
70W
Huawei logo

Da Vinci

2018

datacenter
The 3D Cube matrix unit, a 16x16x16 matrix-multiply block paired with vector and scalar units in one scalable core that Huawei reuses from edge silicon up to datacenter training chips

Da Vinci is Huawei's NPU core architecture, used across the whole Ascend line from edge parts to datacenter training accelerators. Its compute core combines a 3D Cube matrix unit with vector and scalar units, the Cube being the direct architectural analogue of an NVIDIA Tensor Core or an AMD Matrix Core, and every matrix-format throughput figure this database records for an Ascend part is a Cube figure. The first generation is the Ascend 910A, built on TSMC 7nm with 32 GB of HBM2 at 1,228 GB/s and a 310 W envelope. Later generations moved to SMIC after export controls closed TSMC to Huawei. Software is CANN and MindSpore rather than CUDA.

Max HBM
32 GB
Max TDP
310W
AWS logo

NeuronCore-v1

2018

datacenter
The first AWS-designed machine learning core: a power-optimised systolic-array Tensor Engine with FP16, BF16 and INT8 inputs and FP32 or INT32 outputs, beside Vector and Scalar engines and compiler-managed on-chip SRAM

NeuronCore-v1 is the first generation of the AWS Neuron compute core and the engine inside first-generation Inferentia. Each core is an independent heterogeneous compute unit with three engines: a systolic-array Tensor Engine delivering 16 TFLOPS of FP16 or BF16, a Vector Engine at 256 floating point operations per cycle for operations such as layer normalisation and pooling, and a Scalar Engine at 512 operations per cycle for non-linearities such as GELU and sigmoid, all fed from software-managed on-chip SRAM that the compiler uses to keep data local. Four cores make one Inferentia chip, and NeuronLink-v1 lets chips be pipelined so a model can be split across them. NeuronCore-v2, used in Inferentia2 and first-generation Trainium, succeeded it.

Tensor Cores
Gen 1

Vega

2017

both
Rapid Packed Math, which doubled FP16 throughput on the vector units, paired with the first use of HBM2 and a High-Bandwidth Cache Controller in a mainstream Radeon part

Vega is AMD's fifth-generation Graphics Core Next architecture, launched in 2017 and the last GCN design before RDNA split the consumer and datacenter lines apart. Two ideas define it. The first is Rapid Packed Math, which executes two FP16 operations per lane per clock and so doubles half precision throughput against FP32; it is the direct ancestor of every packed-math figure AMD has published since, and it is a vector feature, because Vega has no matrix cores. The second is memory: Vega was the first mainstream Radeon architecture built around HBM2 and a High-Bandwidth Cache Controller, reaching 483.8 GB/s on the 2048-bit Vega 10 die and a full 1,024 GB/s on the 4096-bit 7 nm Vega 20 die used by Radeon VII. Vega is the architecture that spans both sides of AMD's business: the same silicon shipped as the consumer Radeon RX Vega 64 and 56 and Radeon VII, and as the datacenter Radeon Instinct MI25, MI50 and MI60, with the Instinct parts running double precision at half rate where Radeon VII runs it at a quarter. AMD has since dropped every Vega part from ROCm support.

Max HBM
16 GB
Max TDP
300W

Volta

2017

both
The first Tensor Cores: a dedicated mixed-precision matrix-multiply unit alongside the shaders, which is the architectural ancestor of every modern AI accelerator

Volta is NVIDIA's 2017 datacenter architecture and the origin of the Tensor Core. Adding a dedicated mixed-precision matrix unit beside the vector shaders took the V100 from 15.7 TFLOPS FP32 to 125 TFLOPS FP16, an eight-fold step that no shader-only design of the era could approach, and every tensor generation since is a descendant of it. Beyond the TITAN V, the desktop card that put GV100 in a PCIe slot with 12 GB of HBM2 and full-rate FP64, this database holds the V100 in every form NVIDIA sold it: SXM2 and PCIe at 16 GB and 32 GB, plus the later V100S at 32 GB, which lifts memory bandwidth from 900 to 1,134 GB/s and FP16 to 130 TFLOPS. The SXM2 parts carry 300 GB/s of NVLink 2.0 in a 300 W module while the PCIe cards fall back to a 32 GB/s PCIe link at 250 W. FP64 runs at half the FP32 rate, 7.8 against 15.7 TFLOPS, so unlike the later inference-oriented architectures Volta is a genuine HPC part as well as an AI one.

Transistors
21.1B
Max HBM
32 GB
Process
12nm
Max TDP
300W
Tensor Cores
Gen 1

Pascal

2016

both
The first NVIDIA architecture with HBM2 and NVLink, and the first with a packed FP16 path, establishing the template of a high-bandwidth memory-stacked training GPU

Pascal is NVIDIA's 2016 architecture and the point at which the datacenter GPU stopped being a graphics card with ECC. The P100 introduced HBM2 on an interposer, 16 GB at 732 GB/s, and NVLink, 160 GB/s on the SXM2 module, alongside a packed FP16 path that doubles the FP32 rate: 21.2 TFLOPS FP16 against 10.6 TFLOPS FP32 on the SXM2 part, with FP64 at half rate, 5.3 TFLOPS. There are no Tensor Cores in this generation; that arrives with Volta. The Pascal parts here show the split the generation created: the P100 in SXM2 and PCIe forms is the HBM2 training part, while the Tesla P40 (24 GB GDDR5, 250 W) and the passively cooled Tesla P4 (8 GB, 75 W) are GDDR5 inference cards with no HBM and no NVLink. The family's largest capacity, 24 GB, therefore belongs to a GDDR5 card, not to its HBM part.

Transistors
15.3B
Max HBM
16 GB
Process
16nm
Max TDP
300W

Maxwell

2014

both
A ground-up redesign of the SM for perf-per-watt on an unchanged 28 nm node, which is also the last NVIDIA architecture with no half-precision arithmetic of any kind

Maxwell is NVIDIA's 2014 architecture and the generation immediately before the modern AI stack begins. It exists in this database for one reason: it is the baseline. Its SM was rebuilt for performance per watt on the same 28 nm node Kepler used, and the largest die, GM200, reaches 3,072 CUDA cores and 8 billion transistors inside 250 W. What it does not have is the thing every later architecture is sold on. There are no Tensor Cores, and NVIDIA's own CUDA C++ Programming Guide prints the native throughput of 16-bit floating-point add, multiply and multiply-add on compute capability 5.0 and 5.2 as N/A -- Maxwell cannot do half precision at all, at any rate. FP64 runs at 4 results per clock per SM against FP32's 128, a thirty-second, so it is not an HPC architecture either. The one part held here is the GeForce GTX TITAN X of March 2015, the 12 GB GM200 card that was the largest single-GPU frame buffer of its day and the machine a great deal of early deep learning research was actually trained on, in FP32, because there was no alternative.

Transistors
8B
Process
28nm
Max TDP
250W

Side-by-Side Comparison

Compare specs, performance, and innovations across all architectures

Showing 41 architectures
ArchitectureUse CaseProcessTensor Core GenMax HBMKey Innovation
CDNA 6
2027 datacenter 2nm — Gen 6 — — HBM4E memory on a 2nm process, powering the AMD Helios 500 rack-scale platform
CDNA 5 8-die
2026 datacenter TSMC N2 (XCD) / N3 (IOD, FCD) 320B Gen 5 432 GB — Wave32 Work Group Processors, 3D hybrid bonded compute dies, HBM4 and UALink over Ethernet scale-up
Huawei logo
Da Vinci v3
2026 datacenter — — — 144 GB 900W Huawei's first generation with native low-precision compute, adding FP8 and MXFP4 alongside its own HiF8 format, and the first to use Huawei-designed memory, HiBL 1.0 for prefill and HiZQ 2.0 for decode
Qualcomm logo
Dragonfly
2026 datacenter — — — — — 768 GB of LPDDR per accelerator card, the highest per-accelerator memory capacity of any announced AI accelerator, trading memory bandwidth for capacity and cost in rack-scale inference
Maia 200
2026 datacenter TSMC N3 140B — 216 GB 750W Inference-first tile and cluster hierarchy with narrow-precision FP4 and FP8 matrix datapaths, 272MB of software-managed on-die SRAM, and an on-die NIC driving a switchless two-tier Ethernet scale-up fabric
Rubin 2-die
2026 datacenter — 336B — 288 GB — NVFP4 microscaled 4-bit tensor format plus Tensor Core emulated SGEMM and DGEMM, which lift FP32 to 400 TFLOPS and FP64 to 200 TFLOPS on a GPU whose native ALU rates are 130 and 33 TFLOPS
Xe3P
2026 datacenter — — — — 350W LPDDR5X instead of HBM for high-capacity air-cooled inference, with native FP4 and MXFP4 microscaling formats through to FP64
T-Head logo
Zhenwu
2026 datacenter — — — — — ICN, a self-developed inter-chip network that gives every accelerator its own independent ports and memory-semantic addressing into its neighbours, scaled through an ICN Switch to full-bandwidth domains of 64, 128 and 1024 cards
Apple M5
2025 consumer 3nm — — — — —
Blackwell Ultra 2-die
2025 datacenter TSMC 4NP 208B Gen 5 288 GB 1.4KW Enhanced Blackwell with 288GB HBM3e
EF
GCU-CARA 4
2025 datacenter — — — — — A fourth-generation compute unit built on Enflame's own instruction set rather than on a GPGPU model, whose distinguishing feature is native FP8 arithmetic rather than FP8 emulated on wider hardware
RDNA 4
2025 consumer — 53.9B Gen 2 — 304W Second-generation AI accelerators with native FP8 in both E5M2 and E4M3 encodings and 2:1 structured sparsity, taking the matrix rate from twice the vector rate on RDNA 3 to four times it, and making AMD publish a dense and a sparse figure side by side for every matrix precision
MetaX logo
XCORE 1.5
2025 datacenter — — — 144 GB 1.0KW The second iteration of MetaX's compute GPU IP, moving to HBM3e and an OAM module at kilowatt power to compete for frontier training work in the Chinese domestic market
MetaX logo
XCORE 2.5
2025 datacenter — — — — — The next compute instruction set on MetaX's own published ladder, funded and in design but not taped out, with no specification of any kind released
Apple M4
2024 consumer 3nm — — — — —
Battlemage
2024 both 5nm — — — 290W XMX matrix engines widened to 2,048 bits and rebuilt around a native SIMD16 vector engine, putting workstation-class INT8 inference throughput inside a 190 W consumer card and a 70 W half-height one
Blackwell 2-die
2024 both TSMC 4NP 208B Gen 5 192 GB 1.2KW First 2-die GPU design with FP4 precision
RDNA 3.5
2024 consumer 4nm — — — 120W A power-optimised revision of RDNA 3 built for integrated graphics, which on the Strix Halo parts is paired with a 256-bit LPDDR5X memory controller so an integrated GPU finally gets 256 GB/s of bandwidth and up to 96 GB of the system memory addressable as video memory
Tenstorrent logo
Wormhole
2024 both — — — — 300W Ethernet moved onto the die, so a Tensix mesh can be extended across cards and across chassis without a switch or a proprietary fabric, and a two-chip card is programmed as one continuous grid
Apple M3
2023 consumer 3nm — — — — —
Qualcomm logo
Cloud AI 100
2023 datacenter 7nm — — — — 576 MB of on-die SRAM across 64 software-managed AI cores, so that inference weights and activations stay on the die rather than crossing to DRAM
Meta logo
MTIA
2023 datacenter — — — 512 GB 1.7KW High-velocity modular chiplet design: a new accelerator generation roughly every six months, built from reusable compute, network and SoC chiplets, with an inference-first focus and a PyTorch-native software stack
Ada Lovelace
2022 both — — Gen 4 — 450W Fourth-generation Tensor Cores with native FP8, paired with a large L2 cache and Shader Execution Reordering, giving datacenter inference cards Hopper-class low-precision throughput without HBM
Alchemist
2022 both 6nm — Gen 1 — 225W Intel's first generation of discrete gaming GPUs in two decades, the ACM-G10 and ACM-G11 dies implementing Xe-HPG
Apple M2
2022 consumer 5nm — — — — —
Biren logo
Biren SPC
2022 datacenter TSMC 7nm — — 64 GB 550W A chiplet GPGPU presented at Hot Chips 2022 with 2,300 GB/s of BLink scale-up, the most aggressive interconnect any Chinese accelerator had published at the time
Hopper
2022 datacenter TSMC 4N 80B Gen 4 141 GB 700W Transformer Engine with FP8 precision
AWS logo
NeuronCore-v2
2022 datacenter — — Gen 2 32 GB — Configurable FP8 (cFP8) and TF32 datatypes, a programmable GPSIMD engine for custom operators, and hardware support for dynamic shapes and control flow
RDNA 3
2022 consumer 5nm — — — 355W The first chiplet consumer GPU, separating the graphics compute die from memory-cache dies, with dedicated AI accelerators added to the shader array
Ampere
2020 both — 54.2B Gen 3 80 GB 450W Third-generation Tensor Cores introducing TF32 and structured 2:4 sparsity, plus Multi-Instance GPU, which partitions one physical accelerator into up to seven isolated instances
Apple M1
2020 consumer 5nm — — — — —
Graphcore logo
Colossus MK2
2020 datacenter TSMC 7nm 59.4B — — — An accelerator with no DRAM at all, putting 900 MB of compiler-managed SRAM inside the processor beside 1,472 independent MIMD cores, so that every core runs its own program against local memory instead of sharing a scheduler and a cache hierarchy
RDNA 2
2020 consumer 7nm — — — 335W AMD Infinity Cache, a 128 MB on-die last-level cache that let a 256-bit GDDR6 bus feed a 4K-class GPU, alongside the first hardware ray accelerators in a Radeon part
Groq logo
Tensor Streaming Processor
2019 datacenter 14nm 26.8B — — — A processor with no caches, no branch prediction, no reorder buffer and no external memory, where the compiler schedules every instruction issue cycle by cycle, so that execution time is known before the program runs and does not vary between runs
Turing
2019 both 12nm — — — 70W Integer Tensor Core paths, INT8 and INT4, which turned quantised inference into the mainstream deployment format and made a 70 W single-slot accelerator viable
Huawei logo
Da Vinci
2018 datacenter — — — 32 GB 310W The 3D Cube matrix unit, a 16x16x16 matrix-multiply block paired with vector and scalar units in one scalable core that Huawei reuses from edge silicon up to datacenter training chips
AWS logo
NeuronCore-v1
2018 datacenter — — Gen 1 — — The first AWS-designed machine learning core: a power-optimised systolic-array Tensor Engine with FP16, BF16 and INT8 inputs and FP32 or INT32 outputs, beside Vector and Scalar engines and compiler-managed on-chip SRAM
Vega
2017 both — — — 16 GB 300W Rapid Packed Math, which doubled FP16 throughput on the vector units, paired with the first use of HBM2 and a High-Bandwidth Cache Controller in a mainstream Radeon part
Volta
2017 both 12nm 21.1B Gen 1 32 GB 300W The first Tensor Cores: a dedicated mixed-precision matrix-multiply unit alongside the shaders, which is the architectural ancestor of every modern AI accelerator
Pascal
2016 both 16nm 15.3B — 16 GB 300W The first NVIDIA architecture with HBM2 and NVLink, and the first with a packed FP16 path, establishing the template of a high-bandwidth memory-stacked training GPU
Maxwell
2014 both 28nm 8B — — 250W A ground-up redesign of the SM for perf-per-watt on an unchanged 28 nm node, which is also the last NVIDIA architecture with no half-precision arithmetic of any kind

Common Questions

Get answers to frequently asked questions about GPU architectures

Have questions about GPU architectures?

We've answered the most common questions about NVIDIA architectures, Tensor Cores, hardware specs, and performance optimization.

View All FAQs
5 questions answered • Organized by category

Ready to Explore?

Browse our complete AI accelerator database or compare architectures side-by-side

Found this helpful?

Share with your team or bookmark for later

Share Share

© 2026 Flopper.io - Compare the GPUs Powering AI