Removed from NVIDIA's public roadmap at GTC 2026 (March 2026). Announced 9 September 2025 for availability at the end of 2026, but NVIDIA's GTC 2026 Vera Rubin platform materials list six platform chips with no Rubin CPX among them, make no mention of the Vera Rubin NVL144 CPX rack, and the Rubin CPX product page on nvidia.com now returns 404. The part never shipped and no NVIDIA datasheet was ever published, so every figure here is a launch claim. NVIDIA publishes a single NVFP4 figure for this part, 30 petaFLOPS, and does not label it dense or sparse. It is recorded as sparse because NVIDIA's own rack arithmetic requires it: the Vera Rubin NVL144 CPX is stated at 8 exaFLOPS from 144 Rubin CPX and 72 Rubin packages, which reaches 7.92 exaFLOPS only if the Rubin packages contribute their sparse 50 petaFLOPS and the CPX its 30. The dense reading gives 6.84 exaFLOPS. No dense figure is listed because NVIDIA publishes none, and the Rubin line's dense-to-sparse ratio is 1.43 rather than 2, so it cannot be derived.

NVIDIA Rubin CPX

Architecture: RubinAnnounced: 2026Status: Announced, not shippingSpec confidence: Vendor claimed
Memory
128 GB
GDDR7

Overview

NVIDIA Rubin CPX is a context-phase inference accelerator announced on 9 September 2025 at the AI Infrastructure Summit. It is a distinct class of part from the Rubin datacenter GPU: same Rubin architecture, but a cost-efficient monolithic die paired with 128 GB of GDDR7 instead of HBM4, on the reasoning that the prefill or context phase of long-context inference is compute bound rather than memory-bandwidth bound. NVIDIA published only three hardware figures for it: 30 PFLOPS of NVFP4 compute, 128 GB of GDDR7, and 3x the attention acceleration of GB300 NVL72, alongside hardware video encode and decode. Memory bandwidth, TDP, process node, transistor count, form factor, NVLink and PCIe generation were never disclosed and are stored as NULL here rather than estimated. CPX was designed to sit alongside Rubin GPUs in the Vera Rubin NVL144 CPX rack (144 Rubin CPX, 144 Rubin GPUs, 36 Vera CPUs, 8 exaflops NVFP4, 100 TB of memory, 1.7 PB/s of bandwidth), or in a dedicated CPX compute tray for existing Vera Rubin NVL144 installations. Both the chip and that rack disappeared from NVIDIA's roadmap at GTC 2026.

Performance

Peak theoretical throughput by precision type

PrecisionDense2:4 Sparse
FP64
No verified data available Structured sparsity is a tensor-core feature; this vector precision has no sparse form
FP32
No verified data available Structured sparsity is a tensor-core feature; this vector precision has no sparse form
TF32
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
BF16
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
FP16
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
FP8
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
FP6
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
FP4
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision
NVFP4
The vendor publishes only the 2:4 sparse figure for this precision 30,000TFLOPS
INT8
No verified data available The part supports 2:4 sparsity, but the vendor publishes no sparse figure for this precision

Specifications

Architecture

Rubin

Form Factor

No verified data available

Launch Year

2026

Process Node

No verified data available

Memory

128 GB GDDR7

Bandwidth

No verified data available

TDP

No verified data available

Max power

No verified data available

Spec Confidence

Vendor claimed

Full Specifications

Chip Design
Chiplets 1 (Monolithic)
Memory
Memory 128 GB
Memory Type GDDR7
Bandwidth No verified data available
Interface Width No verified data available
Power & Thermal
TDP No verified data available
Max power No verified data available
Enterprise Features
Sparsity Yes
General
Form Factor No verified data available
Architecture Rubin
Process Node No verified data available
Launch Year 2026

Datasheet & Resources

Data Provenance

Every figure traced to a source

Primary Source

Publisher
NVIDIA
Published
2025-09-09

Data Quality

Spec confidence
Vendor claimed
Clock basis
Boost
Core precisions with figures
0 of 9
Normalization
All values in TFLOPS

Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.

Browse GPU Cloud Providers

Similar GPUs

NVIDIA Vera Rubin Superchip

Superchip · 2026
FP32: 260 TFLOPS
Compare vs Vera Rubin

NVIDIA Rubin SXM

SXM · 2026
FP32: 130 TFLOPS
Compare vs Rubin

Frequently Asked Questions

How many TFLOPS does the NVIDIA Rubin CPX have?

The verified peak figures Flopper holds for the NVIDIA Rubin CPX are 30,000 TFLOPS NVFP4 (2:4 sparse). Flopper does not currently have verified FP32, FP16 and FP8 throughput figures for it.

What is the power consumption of the NVIDIA Rubin CPX?

Flopper does not currently have a verified TDP (Thermal Design Power) figure for the NVIDIA Rubin CPX.

How much memory does the NVIDIA Rubin CPX have?

The NVIDIA Rubin CPX is equipped with 128 GB of memory. Flopper does not currently have a verified memory bandwidth figure for it.

What architecture is the NVIDIA Rubin CPX based on?

The NVIDIA Rubin CPX is based on the Rubin architecture, launched in 2026.

Stay Updated on GPU Releases

Get notified when new GPUs are added or specifications are updated.

Loading verification...

No spam, unsubscribe anytime.

Back to GPUs

© 2026 Flopper.io - Compare the GPUs Powering AI