NVIDIA GPU Architectures Explained
Understanding the Evolution from Pascal to Blackwell
NVIDIA's GPU architectures represent major technology platforms that define entire generations of products. Each architecture introduces fundamental innovations in processing capabilities, memory systems, and specialized AI acceleration—spanning from consumer graphics cards to datacenter supercomputers.
Understanding the difference between an architecture (like Hopper), a product (like H100), and a system (like Grace Hopper Superchip) is essential for making informed decisions about GPU infrastructure. This guide breaks down NVIDIA's naming conventions, explains key concepts like Tensor Cores and NVLink, and traces the evolution from Kepler (2012) to today's cutting-edge Blackwell architecture.
What You'll Learn
- How to decode NVIDIA's naming system (Architecture vs. Product vs. Superchip)
- Key innovations in each architecture generation
- Essential concepts: Tensor Cores, multi-die designs, NVLink, and Transformer Engine
- Side-by-side technical comparisons across all datacenter architectures

Decoding NVIDIA's Naming
Architecture vs Product vs Superchip — here's the difference in plain English
Architecture
The blueprint defining a GPU family's core technology and capabilities
Product
A specific GPU model built on an architecture — one arch, many products
Superchip
CPU + GPU unified with ultra-fast interconnects for seamless data sharing
GPU Architecture Evolution Timeline
Click on any architecture to learn more
Essential GPU Concepts
The building blocks that make modern AI acceleration possible
What is a Superchip?
A superchip combines different processor types (CPU + GPU) into one unified system.
NVIDIA Superchips like Grace Hopper integrate an ARM-based Grace CPU with a Hopper GPU using high-bandwidth NVLink-C2C interconnects. This creates a unified memory architecture where CPU and GPU can access the same memory space with 900GB/s bandwidth - 7x faster than PCIe Gen5. This eliminates memory copy overhead and enables seamless data sharing between processors.
What are Tensor Cores?
Specialized hardware units optimized for matrix operations in AI workloads.
Tensor Cores are specialized processing units designed to accelerate matrix multiplication operations fundamental to deep learning. Each generation adds new capabilities: Gen 1 (Volta) introduced FP16, Gen 2 (Turing) added INT8, Gen 3 (Ampere) brought TF32 and structural sparsity, Gen 4 (Hopper/Ada) added FP8 and Transformer Engine, and Gen 5 (Blackwell) introduced FP4 precision. They deliver 10-20x speedup over CUDA cores for AI training and inference.
Multi-Die GPU Architecture
Using multiple silicon dies connected together to create one massive GPU.
Modern chip manufacturing has physical size limits (reticle limit). Multi-die architecture overcomes this by connecting 2+ GPU dies via ultra-fast interconnects, making them appear as a single GPU to software. Blackwell pioneered this with its 2-die design reaching 208B transistors - impossible in a single die. The interconnect bandwidth must be extremely high to avoid bottlenecks between dies.
NVLink: GPU-to-GPU Interconnect
High-speed interconnect for direct GPU-to-GPU communication.
NVLink is NVIDIA's proprietary high-bandwidth, low-latency interconnect for multi-GPU systems. It allows GPUs to directly share memory and communicate without going through the CPU or PCIe bus. NVLink bandwidth has evolved: Gen 1 (Pascal) at 160GB/s, Gen 2 (Volta) at 300GB/s, Gen 3 (Ampere) at 600GB/s, Gen 4 (Hopper) at 900GB/s, and Gen 5 (Blackwell) at 1800GB/s. Essential for training large models across multiple GPUs.
Explore NVIDIA Architectures
Click any architecture to dive deep into specs, peak performance, and GPU models
CDNA 6
2027
CDNA 6 is the sixth generation of AMD's data center compute architecture, named by AMD at CES 2026 and scheduled to arrive with the Instinct MI500 Series in 2027. AMD has stated three things about it and no more: it is built on an advanced 2nm process, it uses HBM4E memory, and it succeeds the CDNA 5 architecture that debuted with the Instinct MI400 series in 2026. AMD has published no compute unit count, no clock speed, no transistor count, no memory capacity or bandwidth, no board power and no throughput figure at any precision for any CDNA 6 part. Its single performance claim is a projection that an MI500-powered AI rack will deliver up to a thousand times the AI performance of an MI300X platform, a comparison AMD makes between a rack and a platform without naming a precision on either side, and which therefore cannot be reduced to a per-accelerator figure. CDNA 6 silicon will power the AMD Helios 500 rack-scale system alongside EPYC "Verano" CPUs and Pensando "Como" and "Monza" networking, and AMD has already named its successor, the MI600 Series in the Helios 600 rack.
CDNA 5
2026
Fifth generation AMD CDNA architecture, introduced with the Instinct MI400 series in 2026. CDNA 5 moves the compute front end to Work Group Processors with Wave32 execution, stacks 3D hybrid bonded compute dies on a CoWoS-L package, and splits the design into eight TSMC N2 Accelerator Complex Dies plus two I/O dies and two Fabric and Cache dies on N3. It is the first Instinct generation on HBM4 and the first to use UALink over Ethernet for rack scale-up, and it quadruples dense OCP MXFP4 and MXFP8 matrix throughput per GPU versus CDNA 4.

Da Vinci v3
2026
Da Vinci v3 is the Ascend 950 generation and the first Huawei architecture with native FP8 and 4-bit compute. It is also the first split into two memory-specialised parts: the Ascend 950PR pairs 128 GB of Huawei's own HiBL 1.0 at 1,600 GB/s for prefill and recommendation work, while the Ascend 950DT pairs 144 GB of HiZQ 2.0 at 4,000 GB/s for decode. Both carry 2,000 GB/s of Unified Bus 2.0 scale-up. The Atlas 350 accelerator card is built on the 950PR chip as a de-rated bin, at seven eighths of its memory and bandwidth. Huawei has published no process node, transistor count or die size for this generation, and the parts are export-controlled to the Chinese domestic market.

Dragonfly
2026
Dragonfly is Qualcomm's brand for its data center AI silicon, introduced with the AI200 and AI250 rack-scale inference accelerators announced on 27 October 2025 and extended by a third-generation AI300 on Qualcomm's published annual cadence. Its defining decision is memory: Dragonfly accelerators carry LPDDR rather than HBM, reaching 768 GB per card where contemporary HBM parts reach 144 GB to 432 GB, and accepting far lower bandwidth in exchange for capacity, cost and supply availability. The design target is large-model inference economics rather than peak throughput, and Qualcomm markets it on total cost of ownership and performance per dollar per watt rather than on FLOPS. Qualcomm builds on its NPU heritage from the Cloud AI 100 family and Snapdragon, but has published no microarchitecture name, no NPU generation number, no AI-core count, no clock speed, no process node and no throughput figure of any kind for any Dragonfly part. The second-generation AI250 adds a near-memory compute architecture that Qualcomm claims delivers more than ten times the effective memory bandwidth of the AI200.
Maia 200
2026
Maia 200 is Microsoft's second-generation in-house AI accelerator architecture and its first silicon and system platform built specifically for inference rather than training. Compute is organised as a hierarchy: a tile pairs a Tile Tensor Unit for matrix multiply with a Tile Vector Processor for programmable SIMD work, backed by multi-banked Tile SRAM, a tile DMA engine and a lightweight Tile Control Processor; tiles compose into clusters that share a large Cluster SRAM and a cluster DMA subsystem staging traffic to co-packaged HBM; clusters compose into the SoC. Both SRAM tiers are fully software managed, so compilers and kernels pin working sets on die rather than relying on cache heuristics, and tile and SRAM redundancy schemes are built in for yield. The Tile Tensor Unit is optimised for FP8, FP6 and FP4 matrix multiplication including mixed FP8 activations against FP4 weights, with hardware casting between storage and compute types at line rate. Scale-up is Ethernet rather than a proprietary fabric: an on-die NIC provides 1.4 TB/s unidirectional bandwidth, Microsoft's AI Transport Layer protocol adds packet spraying, multipath routing and congestion-resistant flow control, groups of four accelerators form a switchless Fully Connected Quad, and a switched second tier extends the domain to 6,144 accelerators.
Rubin
2026
Rubin is NVIDIA's datacenter GPU architecture for 2026, succeeding Blackwell and Blackwell Ultra and named after the astronomer Vera Rubin. A Rubin GPU is a 336 billion transistor part built from two reticle-limited compute dies joined by the NVIDIA High-Bandwidth Interface, carrying 224 streaming multiprocessors and 896 Tensor Cores, paired with 288 GB of HBM4 at 22 TB/s and connected by sixth-generation NVLink at 3.6 TB/s per GPU. It pairs with the 88-core Vera CPU over NVLink-C2C to form the Vera Rubin Superchip, and 72 GPUs with 36 CPUs form the Vera Rubin NVL72 rack. Rubin introduces NVFP4, a microscaled four-bit format with block-level scaling, and a third-generation Transformer Engine with hardware-accelerated adaptive compression. It is also the first NVIDIA architecture to offer both native and Tensor Core emulated FP32 and FP64, so an SGEMM or DGEMM workload sees roughly three and six times the native ALU throughput respectively. The generation also included Rubin CPX, a monolithic GDDR7 part for the context phase of long-context inference, which NVIDIA removed from its roadmap at GTC 2026, and is followed in 2027 by Rubin Ultra.
Xe3P
2026
Xe3P is the performance-optimised variant of Intel's Xe3 graphics architecture. Its first announced datacenter application is Crescent Island, an inference accelerator that pairs Xe3P with LPDDR5X rather than HBM in order to reach very large memory capacity in a 350 W air-cooled PCIe card. Intel has described the datatype range as spanning native FP4 and MXFP4 microscaling formats through to FP64, but has published no throughput, clock, core count or process node figures for any Xe3P part.

Zhenwu
2026
Zhenwu is T-Head's datacenter AI accelerator line, built around a self-developed parallel computing architecture and the ICN inter-chip network, and paired with T-Head SAIL, the company's own software stack. Alibaba describes every member as a 训推一体 part, one chip for both training and inference, and the line is sold as Alibaba Cloud capacity rather than as a component. Three generations exist. The 810E carries 96 GB of HBM2e at 2.7 TB/s with seven ICN ports and 700 GB/s of inter-chip bandwidth. The M890 raises that to 144 GB, eight ports and 800 GB/s, adds FP8 and FP4 to a range that already ran from FP32 through BF16 and FP16, and introduces the ICN Switch, which makes 64 cards look like one flat full-bandwidth domain. The V900 carries 216 GB and 1200 GB/s, deepens the tensor units' FP8 and FP4 instruction precision, adds configurable scaling-factor formats and block sizes for MXFP8 and MXFP4, and scales the same Scale-Up architecture to 1024 chips through two layers of ICN Switch. The defining absence is arithmetic: T-Head has never published a throughput figure for any Zhenwu part, quoting only generational ratios of three times the previous chip.
Apple M5
2025
Blackwell Ultra
2025
Enhanced variant of Blackwell architecture featuring increased HBM capacity (288GB) and improved FP4 performance for the most demanding AI workloads.
GCU-CARA 4
2025
GCU-CARA is Enflame's name for the accelerated compute unit at the centre of its cloud AI chips, and the fourth generation is the one behind the 邃思 400 chip and the L600 module. Enflame's IPO prospectus describes the line as a deliberate refusal of the prevailing design: "公司未跟随英伟达的 GPGPU 架构,而是基于自主指令集", the company did not follow NVIDIA's GPGPU architecture but built on its own instruction set, with GCU-CARA answering the Tensor Core and a separate technology, GCU-LARE, answering NVLink. The instruction set covers compute, control, cache access and data synchronisation, and the microarchitecture has now iterated four times. What the fourth generation adds is native FP8: Enflame describes itself as one of the few Chinese vendors whose silicon supports the datatype in hardware rather than by conversion, which is the same distinction that separates the MetaX C700's FP8 compute from the C600's FP8 conversion instructions. Alongside it sits the fourth-generation GCU-LARE interconnect, which Enflame says supports direct chip-to-chip topologies past the bandwidth and latency limits of PCIe and scales to superpods and clusters of ten thousand cards and beyond. The generation ladder Enflame publishes runs 邃思 1.0 in 2019, 邃思 2.0 and 2.5 in 2021, 邃思 320 in 2024 and 邃思 400 in 2025. Enflame's own filing spells the unit both GCU-CARA, in its product table, and GCU-CARE, in its core-technology section; the table spelling is used here.
RDNA 4
2025
RDNA 4 is AMD's 2025 graphics architecture and the first Radeon generation with AI hardware worth quoting next to a datacenter part. Its second-generation AI accelerators add native FP8 in both E5M2 and E4M3 encodings and 2:1 structured sparsity, neither of which RDNA 3 had. On the 64 compute unit configuration that means 195 TFLOPS of dense FP16 matrix and 389 TFLOPS of dense FP8 on the Radeon RX 9070 XT, four and eight times its 48.7 TFLOPS FP32 vector rate respectively, where RDNA 3 stopped at twice. AMD prints the dense and the structured-sparsity figure side by side on every product page in the family, which makes it one of the easier vendor tables to read correctly. Memory stays GDDR6 rather than HBM, so capacity tops out at 32 GB on the Radeon AI PRO R9700 and bandwidth at 640 GB/s on a 256-bit bus, and that is the binding limit on the family for local model work rather than compute. AMD publishes no FP64 and no BF16 rate for any RDNA 4 part, and no lithography node for any of them either. There is no datacenter RDNA 4 product; Instinct continues on the separate CDNA line.

XCORE 1.5
2025
XCORE 1.5 is the second generation of MetaX's compute GPU IP, named as such in the company's STAR Market filing, and this database holds it as the C600. The part moves the line to HBM3e and an OAM module: 144 GB and up to 1,000 W, with MetaXLink for scale-up. That capacity is a press-corroborated figure sitting inside MetaX's own published floor of more than 96 GB, and it is flagged as such on the GPU row rather than presented as a datasheet value. Memory bandwidth, process node and throughput figures are all unpublished, and a widely circulated set of C600 numbers is disqualified because its power figure contradicts MetaX's own.

XCORE 2.5
2025
XCORE 2.5 is the third generation of MetaX's compute GPU instruction set, named in the company's STAR Market filing as the successor built on the XCORE 1.5 of the C600. This database holds it as the C700, and every numeric column on both rows is empty because MetaX has published nothing: no throughput at any precision, no memory capacity, type or bandwidth, no power figure, no process node and no interconnect. At the date of the filing the project was in physical implementation and the chip had not been manufactured. The widely repeated claim that the C700 approaches an H100 is a qualitative positioning sentence in an IPO document with no number attached, and nothing here derives from it. Two naming traps: XCORE 2.0 is MetaX's RENDER instruction set and is not the compute successor to XCORE 1.5, and MetaX never writes C700 and XCORE 2.5 in the same sentence, so the pairing is an identification from the company's own ISA ladder rather than a MetaX label. The year on both rows is the disclosure year, not a launch year.
Apple M4
2024
Fourth-generation Apple Silicon SoC on second-gen 3nm process. Enhanced GPU performance and power efficiency with LPDDR5X support.
Battlemage
2024
Battlemage is Intel's second generation of discrete Arc graphics, announced on 3 December 2024 and built on the Xe2 microarchitecture at TSMC N5. Its Xe-core is rebuilt around native SIMD16 execution: eight 512-bit vector engines and eight 2,048-bit XMX matrix engines per core, with 256 KB of shared L1 and scratchpad memory, three-way co-issue of floating point, integer and matrix work, and XMX datatype support spanning INT2, INT4, INT8, FP16, BF16 and TF32. Intel claims 70 per cent more performance per Xe-core and 50 per cent more performance per watt than Alchemist. The desktop cards use the BMG-G21 die: five render slices, 20 Xe-cores, 160 XMX engines, 20 ray tracing units, 18 MB of L2 cache and a 192-bit GDDR6 interface. The same die underpins the Arc Pro B50 and B60 workstation cards, which is why this architecture is marked for both consumer and datacenter use; the Arc Pro B60 is sold specifically for local large language model inference and is the building block of Intel's Project Battlematrix workstations, which put eight of them and 192 GB of memory in one Xeon chassis. Intel publishes one throughput figure per Battlemage card, peak INT8 TOPS on XMX with dense models, and no FP32, FP16, BF16, TF32 or FP64 rate for the consumer parts. Note that Xe2 is the microarchitecture and Battlemage the desktop die family, the same relationship Alchemist has to Xe-HPG in this database; no separate Xe2 row exists yet.
Blackwell
2024
Next-generation 2-die GPU architecture with 208B transistors, 5th-gen Tensor Cores supporting FP4 precision, and breakthrough AI inference performance.
RDNA 3.5
2024
RDNA 3.5 is AMD's integrated-graphics revision of RDNA 3, announced at Computex on 2 June 2024 alongside the Ryzen AI 300 series and the Radeon 800M graphics that ship inside it. It is not a discrete generation: there is no RDNA 3.5 graphics card, and AMD publishes no compute-unit-level architectural detail for it in the way it does for RDNA 3 and RDNA 4. What makes it matter for this catalog is the Strix Halo configuration, where AMD pairs up to 40 RDNA 3.5 compute units with a 256-bit LPDDR5X memory controller. That gives an integrated GPU 256 GB/s of memory bandwidth and, through AMD's Variable Graphics Memory, up to 96 GB of a 128 GB unified pool addressable as video memory, which is a great deal more capacity than any consumer discrete card offers and the reason these parts get bought for running large models locally. AMD publishes no FP32, FP16, BF16 or INT8 throughput figure for any RDNA 3.5 integrated GPU, and no AI Accelerator or Ray Accelerator count either, so the only compute number available for these parts is a vector rate computed on AMD's own arithmetic. The TOPS figures AMD does publish for Ryzen AI describe the separate XDNA 2 neural processor or the whole SoC, not the graphics.

Wormhole
2024
Wormhole is the second generation of Tenstorrent's Tensix architecture, following Grayskull and preceding Blackhole. A Tensix core is not a shader: it is a matrix engine and a SIMD vector unit wrapped around five small RISC-V cores that handle unpack, math, pack and two network-on-chip interfaces, with roughly 1.5 MB of local SRAM each. The defining change in Wormhole is that Ethernet is on the die, so the mesh of Tensix cores extends off-chip over standard Ethernet rather than a proprietary fabric, which is how a two-processor n300 card presents as one continuous grid and how cards are meshed to each other through their QSFP-DD ports without going through the host. Tenstorrent publishes three compute figures for Wormhole, FP8, FP16 and BFP8, and marks none of them sparse, because the hardware has no structured sparsity. It also lists a long tail of supported formats it publishes no rate for, including BF16, TF32, INT8, INT32, several block-float widths and its own VTF19. The whole software stack, tt-metal and tt-forge, is open source, which together with a card price near a thousand dollars is most of the reason these parts get bought.
Apple M3
2023
Third-generation Apple Silicon SoC built on TSMC 3nm process. Features Dynamic Caching, hardware-accelerated ray tracing, and mesh shading.

Cloud AI 100
2023
Cloud AI 100 is Qualcomm's first data center AI architecture, a fixed-function inference design built on a 7 nm process and sold as PCIe cards rather than as a rack-scale system. Its defining choice is on-die memory: each AI core owns 9 MB of software-managed SRAM, giving 576 MB on the largest card, which is more local scratchpad memory than any other non-wafer-scale accelerator carries and which lets inference workloads avoid DRAM traffic entirely for many layers. Capacity comes from LPDDR4x rather than HBM, trading bandwidth for cost and power in a 150 W envelope. Qualcomm publishes throughput for exactly two datatypes, INT8 and FP16, and no architectural detail beyond the AI core count and its SRAM allocation: there is no published clock speed, no MAC width, no die size and no transistor count for any part in the family. The line spans five SKUs cut from the same core, differing in how many cores are enabled and at what clock, from a 16-core 75 W card to a 64-core 150 W card. Qualcomm has never applied its later Dragonfly data center brand to this family, and never uses the word Hexagon in any Cloud AI 100 document. It remains the only Qualcomm data center accelerator that has shipped and can be rented, its Dragonfly successors having been announced with no published throughput figures and no availability.

MTIA
2023
MTIA, the Meta Training and Inference Accelerator, is Meta's family of in-house AI accelerators, developed in close partnership with Broadcom and deployed only inside Meta's own data centers. The architecture is organised around a grid of processing elements. Each PE contains two RISC-V vector cores, a Dot Product Engine for matrix multiplication, a Special Function Unit for activations and elementwise operations, a Reduction Engine for accumulation and inter-PE communication, and a DMA engine feeding local scratch memory. Chips are assembled from modular chiplets, which is what lets Meta ship a new generation on a roughly six-month cadence rather than a two-year one: MTIA 300 uses one compute chiplet with two network chiplets, MTIA 400 combines two compute chiplets to double compute density, and MTIA 500 moves to a two-by-two arrangement of smaller compute chiplets surrounded by HBM stacks, two network chiplets and an SoC chiplet providing PCIe connectivity to the host. The design philosophy is explicitly inference-first and cost-driven rather than peak-FLOPS driven, with built-in NIC chiplets, dedicated message engines that offload communication collectives, and near-memory compute for reduction-based collectives. Later generations add hardware acceleration for attention and mixture-of-experts bottlenecks and custom low-precision data types intended to raise throughput without hurting model quality. Across the generation Meta states that HBM bandwidth rises 4.5 times and compute rises 25 times, though that 25x figure compares MTIA 300 at 8-bit precision with MTIA 500 at 4-bit precision rather than measuring like for like. Software is PyTorch native, with a graph compiler, a Triton compiler and eager-mode runtime support, which Meta treats as central to getting internal workloads onto the hardware quickly.
Ada Lovelace
2022
Ada Lovelace is NVIDIA's 2022 graphics architecture and the generation that made FP8 inference cheap on GDDR6 cards. Its fourth-generation Tensor Cores add native FP8 support, and this database holds Ada parts across three market segments: the datacenter L-series (L40S, L40, L20, L4, L2), the RTX workstation cards (RTX 6000 Ada, RTX 5000 Ada, RTX 2000 Ada) and GeForce (RTX 4090, 4080, 4080 Super, 4070 Ti). The L40S is the family's inference workhorse at 48 GB of GDDR6, 864 GB/s, 91.6 TFLOPS FP32 and 733 TFLOPS dense FP8 in a 350 W PCIe card. The trade-off against a datacenter architecture is FP64: the L40S manages 1.45 TFLOPS against its own 91.6 TFLOPS FP32, roughly a sixty-fourth, so Ada is an inference and visualisation architecture rather than an HPC one. It also has no NVLink on any part in this database, and no HBM anywhere in the family.
Alchemist
2022
Alchemist is the codename for Intel's first-generation Arc discrete GPU family, built on the Xe-HPG microarchitecture across two dies, ACM-G10 and ACM-G11. This database holds the Arc A770 with 16 GB of GDDR6 at 560 GB/s and a 225 W board power. Alchemist is recorded as a child of Xe-HPG: Xe-HPG is the microarchitecture and Alchemist is the die family that implements it, the same relationship the Data Center GPU Flex Series parts have to Xe-HPG.
Apple M2
2022
Second-generation Apple Silicon SoC. Up to 10-core GPU with improved performance per watt and higher memory bandwidth.

Biren SPC
2022
Biren SPC is Biren Technology's first-generation general-purpose GPU architecture. The name is a label used by this database rather than a Biren one: the company calls it simply its first-generation GPGPU architecture and has never given it a name. The specified part is the BR100, presented at Hot Chips 2022 and for a while the most ambitious Chinese datacenter GPU design published anywhere: a 64 GB HBM2e OAM module at 1,600 GB/s and 550 W, rated at 256 TFLOPS FP32, 512 TFLOPS TF32, 1,024 TFLOPS BF16 and 2,048 TOPS INT8, with 2,300 GB/s of BLink scale-up bandwidth. Those figures come from Biren's own conference presentation rather than a datasheet, so they are recorded as a vendor claim. The inference-oriented BR110 shares this architecture and is also held here, but with no specifications at all: Biren has never published a single throughput, memory, power or process figure for it, and nothing on that row is inherited from the BR100. Read the generation in its commercial context, which is the largest fact about it: Biren has been under US export restrictions since October 2023, which cut it off from leading-edge foundry capacity, so shipped volume is very much smaller than the specification implies.
Hopper
2022
Fourth-generation datacenter GPU architecture featuring Transformer Engine, 4th-gen Tensor Cores with FP8 support, and enhanced NVLink for AI training and HPC workloads.

NeuronCore-v2
2022
NeuronCore-v2 is the second generation of the AWS Neuron compute core. Each core is a fully independent heterogeneous compute unit with four engines: a systolic-array Tensor Engine delivering over 90 TFLOPS of FP16 or BF16, a Vector Engine at 2.3 TFLOPS of FP32, a Scalar Engine at 2.9 TFLOPS of FP32, and a new GPSIMD Engine of eight fully programmable 512-bit vector processors that run general-purpose C code against the embedded on-chip SRAM, which is how custom operators are implemented. The generation introduces the configurable FP8 datatype AWS calls cFP8 alongside TF32, adds control flow, dynamic shapes and programmable rounding including stochastic rounding, and scales out over NeuronLink-v2. Two NeuronCore-v2 cores make up one Inferentia2 chip.
RDNA 3
2022
RDNA 3 is the graphics half of AMD's split roadmap, the consumer counterpart to CDNA 3, and the first consumer GPU architecture built from chiplets: a graphics compute die packaged with separate memory-cache dies. This database holds two RDNA 3 parts, the Radeon RX 7900 XTX (24 GB, 960 GB/s, 355 W) and the RX 7900 XT (20 GB, 800 GB/s, 315 W). They are included for cross-architecture comparison rather than as datacenter options, and the numbers show why: the 7900 XTX delivers 61.4 TFLOPS FP32 and 122.8 TFLOPS FP16 through its AI accelerators, a 2x step where a CDNA or Tensor Core part manages 8x or more, and its FP64 rate is 1.54 TFLOPS, about a fortieth of FP32. There is no HBM, no Infinity Fabric scale-up and no ECC HBM stack anywhere in the family.
Ampere
2020
Ampere is NVIDIA's 2020 architecture and the widest family in this database, spanning the datacenter GA100 die and the consumer and professional GA10x dies. Its third-generation Tensor Cores introduced TF32, a 19-bit format that runs unmodified FP32 training code on the tensor path, and structured 2:4 sparsity, which doubles the quoted rate when a model is pruned to that pattern -- every value stored here is the DENSE figure. The A100 is the reference part: 80 GB of HBM2e at 2,039 GB/s in the SXM4 form factor, 19.5 TFLOPS FP32, 156 TFLOPS TF32, 312 TFLOPS FP16 and BF16, 624 TOPS INT8, and 9.7 TFLOPS FP64 at half the FP32 rate, with 600 GB/s of NVLink. Ampere also brought Multi-Instance GPU, and it is the generation the export-controlled A800 belongs to. Two nodes are used across the family, 7 nm for GA100 and 8 nm for GA10x, so no single process node is true of the architecture.
Apple M1
2020
First-generation Apple Silicon SoC. 8-core GPU based on Apple's custom GPU architecture with unified memory.

Colossus MK2
2020
Colossus MK2 is the second generation of Graphcore's Intelligence Processing Unit, announced in July 2020 as the GC200. Each chip is 59.4 billion transistors on an 823 square millimetre TSMC 7nm die, organised as 1,472 IPU-Cores running 8,832 worker threads, and its defining property is that all of its working memory, 900 MB of it, is SRAM inside the processor. There is no HBM, no GDDR and no cache hierarchy: memory is addressed directly and scheduled by the Poplar compiler, and the tile-local arrangement is what gives the architecture its very high on-chip bandwidth. Cores are fully independent, which suits sparse and irregular graph work that a lockstep SIMT machine handles badly, and the trade is that a model larger than 900 MB has to be split across chips over the IPU-Link fabric or streamed from DDR4 on the host machine. Numerics follow Graphcore's AI-Float convention, which names the multiply and accumulate widths separately as FP16.16 and FP16.32, with stochastic rounding in hardware so that FP16 master weights hold their accuracy. Three parts share the generation. The GC200 is the original, rated at 250 TFLOPS of FP16.16 and 62.5 TFLOPS of FP32. The Bow IPU of March 2022 is the same architecture built as a wafer-on-wafer 3D stack, with a power delivery die bonded underneath that lets the logic clock at 1.85 GHz for 350 TFLOPS. The C600 of November 2022 is a revision with FP8 added, the only member sold as a PCIe card, reaching 560 TFLOPS of FP8 and 280 of FP16 at 185 W.
RDNA 2
2020
RDNA 2 is AMD's 2020 graphics architecture, the generation that powers the PlayStation 5 and Xbox Series consoles as well as the Radeon RX 6000 and Radeon PRO W6000 desktop cards. Its defining idea is Infinity Cache: up to 128 MB of on-die last-level cache that absorbs enough traffic to let a 256-bit GDDR6 bus feed an 80 compute unit GPU, which is how the Radeon RX 6900 XT reaches 23 TFLOPS of FP32 on 512 GB/s of external bandwidth. It also brought the first hardware ray accelerators to Radeon, one per compute unit. For machine learning the important characteristic is a negative one: RDNA 2 has no matrix cores at all. Every performance figure AMD publishes for it is a vector figure, with FP16 running at exactly twice FP32 through packed math, and there is no BF16, FP8, INT8 or INT4 throughput published anywhere in the generation. Matrix hardware arrived on the Radeon side one generation later, with RDNA 3. AMD built RDNA 2 on TSMC 7 nm across three dies, Navi 21 at 26.8 billion transistors, Navi 22 at 17.2 billion and Navi 23 at 11.1 billion, and every part in the family uses GDDR6 rather than HBM.

Tensor Streaming Processor
2019
The Tensor Streaming Processor is the architecture Groq announced in November 2019 and later rebranded the LPU. It inverts the conventional layout: instead of a two-dimensional mesh of identical cores each holding a slice of every function, the TSP slices the functions themselves into parallel lanes that run the width of the die, so that data streams sideways through a memory slice, a vector slice, a matrix slice and a switch slice in turn. Nothing is scheduled by hardware. There are no caches to miss, no branch prediction, no reorder buffer and no external memory of any kind, only 230 MB of globally shared SRAM addressed directly at up to 80 TB/s, with the compiler placing every instruction on a known cycle. The result is determinism: the same program takes exactly the same number of cycles every time it runs, which is why Groq sells the design on tail latency rather than on peak throughput. The matrix unit is an integer array, and floating point arrives through Groq's TruePoint technique, which decomposes an FP16 or FP32 multiply into integer work on the same hardware; that is why the first-generation part is rated four times faster at INT8 than at FP16 rather than twice. The trade is capacity. A model larger than 230 MB has to be split across many chips, so the architecture spends its area on very high radix chip-to-chip links instead of on local memory. First silicon was a 25 by 29 millimetre 14nm die of 26.8 billion transistors running at a nominal 900 MHz.
Turing
2019
Turing is NVIDIA's 2018-2019 architecture, and this database holds it across the datacenter T4, the Quadro RTX workstation pair and the GeForce and Titan cards that introduced ray tracing. Its contribution to inference was integer precision: adding INT8 and INT4 paths to the Tensor Core took the T4 to 130 TOPS INT8 and 260 TOPS INT4 against 65 TFLOPS FP16 and 8.1 TFLOPS FP32, and it did so inside a 70 W single-slot passively cooled PCIe card with 16 GB of GDDR6 at 300 GB/s. That power envelope is the point: the T4 became the default inference card in dense server deployments for years. There is no NVLink on the T4 and no HBM anywhere in the family, and the 32 GB/s recorded on its row is its PCIe link, not a scale-up fabric. Turing is also the generation that introduced hardware ray tracing on the consumer side, and the GeForce RTX 20 series and TITAN RTX rows here carry that consumer side of the family.

Da Vinci
2018
Da Vinci is Huawei's NPU core architecture, used across the whole Ascend line from edge parts to datacenter training accelerators. Its compute core combines a 3D Cube matrix unit with vector and scalar units, the Cube being the direct architectural analogue of an NVIDIA Tensor Core or an AMD Matrix Core, and every matrix-format throughput figure this database records for an Ascend part is a Cube figure. The first generation is the Ascend 910A, built on TSMC 7nm with 32 GB of HBM2 at 1,228 GB/s and a 310 W envelope. Later generations moved to SMIC after export controls closed TSMC to Huawei. Software is CANN and MindSpore rather than CUDA.

NeuronCore-v1
2018
NeuronCore-v1 is the first generation of the AWS Neuron compute core and the engine inside first-generation Inferentia. Each core is an independent heterogeneous compute unit with three engines: a systolic-array Tensor Engine delivering 16 TFLOPS of FP16 or BF16, a Vector Engine at 256 floating point operations per cycle for operations such as layer normalisation and pooling, and a Scalar Engine at 512 operations per cycle for non-linearities such as GELU and sigmoid, all fed from software-managed on-chip SRAM that the compiler uses to keep data local. Four cores make one Inferentia chip, and NeuronLink-v1 lets chips be pipelined so a model can be split across them. NeuronCore-v2, used in Inferentia2 and first-generation Trainium, succeeded it.
Vega
2017
Vega is AMD's fifth-generation Graphics Core Next architecture, launched in 2017 and the last GCN design before RDNA split the consumer and datacenter lines apart. Two ideas define it. The first is Rapid Packed Math, which executes two FP16 operations per lane per clock and so doubles half precision throughput against FP32; it is the direct ancestor of every packed-math figure AMD has published since, and it is a vector feature, because Vega has no matrix cores. The second is memory: Vega was the first mainstream Radeon architecture built around HBM2 and a High-Bandwidth Cache Controller, reaching 483.8 GB/s on the 2048-bit Vega 10 die and a full 1,024 GB/s on the 4096-bit 7 nm Vega 20 die used by Radeon VII. Vega is the architecture that spans both sides of AMD's business: the same silicon shipped as the consumer Radeon RX Vega 64 and 56 and Radeon VII, and as the datacenter Radeon Instinct MI25, MI50 and MI60, with the Instinct parts running double precision at half rate where Radeon VII runs it at a quarter. AMD has since dropped every Vega part from ROCm support.
Volta
2017
Volta is NVIDIA's 2017 datacenter architecture and the origin of the Tensor Core. Adding a dedicated mixed-precision matrix unit beside the vector shaders took the V100 from 15.7 TFLOPS FP32 to 125 TFLOPS FP16, an eight-fold step that no shader-only design of the era could approach, and every tensor generation since is a descendant of it. Beyond the TITAN V, the desktop card that put GV100 in a PCIe slot with 12 GB of HBM2 and full-rate FP64, this database holds the V100 in every form NVIDIA sold it: SXM2 and PCIe at 16 GB and 32 GB, plus the later V100S at 32 GB, which lifts memory bandwidth from 900 to 1,134 GB/s and FP16 to 130 TFLOPS. The SXM2 parts carry 300 GB/s of NVLink 2.0 in a 300 W module while the PCIe cards fall back to a 32 GB/s PCIe link at 250 W. FP64 runs at half the FP32 rate, 7.8 against 15.7 TFLOPS, so unlike the later inference-oriented architectures Volta is a genuine HPC part as well as an AI one.
Pascal
2016
Pascal is NVIDIA's 2016 architecture and the point at which the datacenter GPU stopped being a graphics card with ECC. The P100 introduced HBM2 on an interposer, 16 GB at 732 GB/s, and NVLink, 160 GB/s on the SXM2 module, alongside a packed FP16 path that doubles the FP32 rate: 21.2 TFLOPS FP16 against 10.6 TFLOPS FP32 on the SXM2 part, with FP64 at half rate, 5.3 TFLOPS. There are no Tensor Cores in this generation; that arrives with Volta. The Pascal parts here show the split the generation created: the P100 in SXM2 and PCIe forms is the HBM2 training part, while the Tesla P40 (24 GB GDDR5, 250 W) and the passively cooled Tesla P4 (8 GB, 75 W) are GDDR5 inference cards with no HBM and no NVLink. The family's largest capacity, 24 GB, therefore belongs to a GDDR5 card, not to its HBM part.
Maxwell
2014
Maxwell is NVIDIA's 2014 architecture and the generation immediately before the modern AI stack begins. It exists in this database for one reason: it is the baseline. Its SM was rebuilt for performance per watt on the same 28 nm node Kepler used, and the largest die, GM200, reaches 3,072 CUDA cores and 8 billion transistors inside 250 W. What it does not have is the thing every later architecture is sold on. There are no Tensor Cores, and NVIDIA's own CUDA C++ Programming Guide prints the native throughput of 16-bit floating-point add, multiply and multiply-add on compute capability 5.0 and 5.2 as N/A -- Maxwell cannot do half precision at all, at any rate. FP64 runs at 4 results per clock per SM against FP32's 128, a thirty-second, so it is not an HPC architecture either. The one part held here is the GeForce GTX TITAN X of March 2015, the 12 GB GM200 card that was the largest single-GPU frame buffer of its day and the machine a great deal of early deep learning research was actually trained on, in FP32, because there was no alternative.
Side-by-Side Comparison
Compare specs, performance, and innovations across all architectures
| Architecture | Use Case | Process | Tensor Core Gen | Max HBM | Key Innovation | |||
|---|---|---|---|---|---|---|---|---|
| CDNA 6 | 2027 | datacenter | 2nm | — | Gen 6 | — | — | HBM4E memory on a 2nm process, powering the AMD Helios 500 rack-scale platform |
| CDNA 5 8-die | 2026 | datacenter | TSMC N2 (XCD) / N3 (IOD, FCD) | 320B | Gen 5 | 432 GB | — | Wave32 Work Group Processors, 3D hybrid bonded compute dies, HBM4 and UALink over Ethernet scale-up |
![]() | 2026 | datacenter | — | — | — | 144 GB | 900W | Huawei's first generation with native low-precision compute, adding FP8 and MXFP4 alongside its own HiF8 format, and the first to use Huawei-designed memory, HiBL 1.0 for prefill and HiZQ 2.0 for decode |
![]() | 2026 | datacenter | — | — | — | — | — | 768 GB of LPDDR per accelerator card, the highest per-accelerator memory capacity of any announced AI accelerator, trading memory bandwidth for capacity and cost in rack-scale inference |
| Maia 200 | 2026 | datacenter | TSMC N3 | 140B | — | 216 GB | 750W | Inference-first tile and cluster hierarchy with narrow-precision FP4 and FP8 matrix datapaths, 272MB of software-managed on-die SRAM, and an on-die NIC driving a switchless two-tier Ethernet scale-up fabric |
| Rubin 2-die | 2026 | datacenter | — | 336B | — | 288 GB | — | NVFP4 microscaled 4-bit tensor format plus Tensor Core emulated SGEMM and DGEMM, which lift FP32 to 400 TFLOPS and FP64 to 200 TFLOPS on a GPU whose native ALU rates are 130 and 33 TFLOPS |
| Xe3P | 2026 | datacenter | — | — | — | — | 350W | LPDDR5X instead of HBM for high-capacity air-cooled inference, with native FP4 and MXFP4 microscaling formats through to FP64 |
![]() | 2026 | datacenter | — | — | — | — | — | ICN, a self-developed inter-chip network that gives every accelerator its own independent ports and memory-semantic addressing into its neighbours, scaled through an ICN Switch to full-bandwidth domains of 64, 128 and 1024 cards |
| Apple M5 | 2025 | consumer | 3nm | — | — | — | — | — |
| Blackwell Ultra 2-die | 2025 | datacenter | TSMC 4NP | 208B | Gen 5 | 288 GB | 1.4KW | Enhanced Blackwell with 288GB HBM3e |
EF | 2025 | datacenter | — | — | — | — | — | A fourth-generation compute unit built on Enflame's own instruction set rather than on a GPGPU model, whose distinguishing feature is native FP8 arithmetic rather than FP8 emulated on wider hardware |
| RDNA 4 | 2025 | consumer | — | 53.9B | Gen 2 | — | 304W | Second-generation AI accelerators with native FP8 in both E5M2 and E4M3 encodings and 2:1 structured sparsity, taking the matrix rate from twice the vector rate on RDNA 3 to four times it, and making AMD publish a dense and a sparse figure side by side for every matrix precision |
![]() | 2025 | datacenter | — | — | — | 144 GB | 1.0KW | The second iteration of MetaX's compute GPU IP, moving to HBM3e and an OAM module at kilowatt power to compete for frontier training work in the Chinese domestic market |
![]() | 2025 | datacenter | — | — | — | — | — | The next compute instruction set on MetaX's own published ladder, funded and in design but not taped out, with no specification of any kind released |
| Apple M4 | 2024 | consumer | 3nm | — | — | — | — | — |
| Battlemage | 2024 | both | 5nm | — | — | — | 290W | XMX matrix engines widened to 2,048 bits and rebuilt around a native SIMD16 vector engine, putting workstation-class INT8 inference throughput inside a 190 W consumer card and a 70 W half-height one |
| Blackwell 2-die | 2024 | both | TSMC 4NP | 208B | Gen 5 | 192 GB | 1.2KW | First 2-die GPU design with FP4 precision |
| RDNA 3.5 | 2024 | consumer | 4nm | — | — | — | 120W | A power-optimised revision of RDNA 3 built for integrated graphics, which on the Strix Halo parts is paired with a 256-bit LPDDR5X memory controller so an integrated GPU finally gets 256 GB/s of bandwidth and up to 96 GB of the system memory addressable as video memory |
![]() | 2024 | both | — | — | — | — | 300W | Ethernet moved onto the die, so a Tensix mesh can be extended across cards and across chassis without a switch or a proprietary fabric, and a two-chip card is programmed as one continuous grid |
| Apple M3 | 2023 | consumer | 3nm | — | — | — | — | — |
![]() | 2023 | datacenter | 7nm | — | — | — | — | 576 MB of on-die SRAM across 64 software-managed AI cores, so that inference weights and activations stay on the die rather than crossing to DRAM |
![]() | 2023 | datacenter | — | — | — | 512 GB | 1.7KW | High-velocity modular chiplet design: a new accelerator generation roughly every six months, built from reusable compute, network and SoC chiplets, with an inference-first focus and a PyTorch-native software stack |
| Ada Lovelace | 2022 | both | — | — | Gen 4 | — | 450W | Fourth-generation Tensor Cores with native FP8, paired with a large L2 cache and Shader Execution Reordering, giving datacenter inference cards Hopper-class low-precision throughput without HBM |
| Alchemist | 2022 | both | 6nm | — | Gen 1 | — | 225W | Intel's first generation of discrete gaming GPUs in two decades, the ACM-G10 and ACM-G11 dies implementing Xe-HPG |
| Apple M2 | 2022 | consumer | 5nm | — | — | — | — | — |
![]() | 2022 | datacenter | TSMC 7nm | — | — | 64 GB | 550W | A chiplet GPGPU presented at Hot Chips 2022 with 2,300 GB/s of BLink scale-up, the most aggressive interconnect any Chinese accelerator had published at the time |
| Hopper | 2022 | datacenter | TSMC 4N | 80B | Gen 4 | 141 GB | 700W | Transformer Engine with FP8 precision |
![]() | 2022 | datacenter | — | — | Gen 2 | 32 GB | — | Configurable FP8 (cFP8) and TF32 datatypes, a programmable GPSIMD engine for custom operators, and hardware support for dynamic shapes and control flow |
| RDNA 3 | 2022 | consumer | 5nm | — | — | — | 355W | The first chiplet consumer GPU, separating the graphics compute die from memory-cache dies, with dedicated AI accelerators added to the shader array |
| Ampere | 2020 | both | — | 54.2B | Gen 3 | 80 GB | 450W | Third-generation Tensor Cores introducing TF32 and structured 2:4 sparsity, plus Multi-Instance GPU, which partitions one physical accelerator into up to seven isolated instances |
| Apple M1 | 2020 | consumer | 5nm | — | — | — | — | — |
![]() | 2020 | datacenter | TSMC 7nm | 59.4B | — | — | — | An accelerator with no DRAM at all, putting 900 MB of compiler-managed SRAM inside the processor beside 1,472 independent MIMD cores, so that every core runs its own program against local memory instead of sharing a scheduler and a cache hierarchy |
| RDNA 2 | 2020 | consumer | 7nm | — | — | — | 335W | AMD Infinity Cache, a 128 MB on-die last-level cache that let a 256-bit GDDR6 bus feed a 4K-class GPU, alongside the first hardware ray accelerators in a Radeon part |
![]() | 2019 | datacenter | 14nm | 26.8B | — | — | — | A processor with no caches, no branch prediction, no reorder buffer and no external memory, where the compiler schedules every instruction issue cycle by cycle, so that execution time is known before the program runs and does not vary between runs |
| Turing | 2019 | both | 12nm | — | — | — | 70W | Integer Tensor Core paths, INT8 and INT4, which turned quantised inference into the mainstream deployment format and made a 70 W single-slot accelerator viable |
![]() | 2018 | datacenter | — | — | — | 32 GB | 310W | The 3D Cube matrix unit, a 16x16x16 matrix-multiply block paired with vector and scalar units in one scalable core that Huawei reuses from edge silicon up to datacenter training chips |
![]() | 2018 | datacenter | — | — | Gen 1 | — | — | The first AWS-designed machine learning core: a power-optimised systolic-array Tensor Engine with FP16, BF16 and INT8 inputs and FP32 or INT32 outputs, beside Vector and Scalar engines and compiler-managed on-chip SRAM |
| Vega | 2017 | both | — | — | — | 16 GB | 300W | Rapid Packed Math, which doubled FP16 throughput on the vector units, paired with the first use of HBM2 and a High-Bandwidth Cache Controller in a mainstream Radeon part |
| Volta | 2017 | both | 12nm | 21.1B | Gen 1 | 32 GB | 300W | The first Tensor Cores: a dedicated mixed-precision matrix-multiply unit alongside the shaders, which is the architectural ancestor of every modern AI accelerator |
| Pascal | 2016 | both | 16nm | 15.3B | — | 16 GB | 300W | The first NVIDIA architecture with HBM2 and NVLink, and the first with a packed FP16 path, establishing the template of a high-bandwidth memory-stacked training GPU |
| Maxwell | 2014 | both | 28nm | 8B | — | — | 250W | A ground-up redesign of the SM for perf-per-watt on an unchanged 28 nm node, which is also the last NVIDIA architecture with no half-precision arithmetic of any kind |
Common Questions
Get answers to frequently asked questions about GPU architectures
Have questions about GPU architectures?
We've answered the most common questions about NVIDIA architectures, Tensor Cores, hardware specs, and performance optimization.
View All FAQsReady to Explore?
Browse our complete AI accelerator database or compare architectures side-by-side