Removed from NVIDIA's public roadmap at GTC 2026 (March 2026). Announced 9 September 2025 for availability at the end of 2026, but NVIDIA's GTC 2026 Vera Rubin platform materials list six platform chips with no Rubin CPX among them, make no mention of the Vera Rubin NVL144 CPX rack, and the Rubin CPX product page on nvidia.com now returns 404. The part never shipped and no NVIDIA datasheet was ever published, so every figure here is a launch claim. NVIDIA publishes a single NVFP4 figure for this part, 30 petaFLOPS, and does not label it dense or sparse. It is recorded as sparse because NVIDIA's own rack arithmetic requires it: the Vera Rubin NVL144 CPX is stated at 8 exaFLOPS from 144 Rubin CPX and 72 Rubin packages, which reaches 7.92 exaFLOPS only if the Rubin packages contribute their sparse 50 petaFLOPS and the CPX its 30. The dense reading gives 6.84 exaFLOPS. No dense figure is listed because NVIDIA publishes none, and the Rubin line's dense-to-sparse ratio is 1.43 rather than 2, so it cannot be derived.

NVIDIA Rubin CPX NEW

Rubin 2026
VRAM
128 GB
GDDR7

Overview

NVIDIA Rubin CPX is a context-phase inference accelerator announced on 9 September 2025 at the AI Infrastructure Summit. It is a distinct class of part from the Rubin datacenter GPU: same Rubin architecture, but a cost-efficient monolithic die paired with 128 GB of GDDR7 instead of HBM4, on the reasoning that the prefill or context phase of long-context inference is compute bound rather than memory-bandwidth bound. NVIDIA published only three hardware figures for it: 30 PFLOPS of NVFP4 compute, 128 GB of GDDR7, and 3x the attention acceleration of GB300 NVL72, alongside hardware video encode and decode. Memory bandwidth, TDP, process node, transistor count, form factor, NVLink and PCIe generation were never disclosed and are stored as NULL here rather than estimated. CPX was designed to sit alongside Rubin GPUs in the Vera Rubin NVL144 CPX rack (144 Rubin CPX, 144 Rubin GPUs, 36 Vera CPUs, 8 exaflops NVFP4, 100 TB of memory, 1.7 PB/s of bandwidth), or in a dedicated CPX compute tray for existing Vera Rubin NVL144 installations. Both the chip and that rack disappeared from NVIDIA's roadmap at GTC 2026.

Performance Metrics

Peak theoretical throughput by precision type

PrecisionBitsPeak TFLOPS
NVFP4 -- 0.0

Power Specifications

TDP

--

Max Power

--

Power Connector

PCIe Slot

Cooling

Air

Memory Specifications

Capacity

128 GB

Type

GDDR7

Bandwidth

--

Interface

--

Hardware & Design

Form Factor

--

Architecture

Rubin

Process Node

--

Launch Year

2026

Variant

Standard

Market Segment

Professional

Chiplets

1 (Monolithic)

Full Specifications

Chip Design
Chiplets 1 (Monolithic)
Memory
VRAM 128 GB
Memory Type GDDR7
Power & Thermal
Enterprise Features
Sparsity (2:4) Yes
General
Architecture Rubin
Launch Year 2026

Documentation & Resources

Common Use Cases

General Compute AI/ML Workloads Data Processing

The NVIDIA Rubin CPX is optimized for high-performance computing tasks with Rubin architecture delivering high TFLOPS of compute power.

Where to Rent

Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.

Browse GPU Cloud Providers

Similar GPUs

NVIDIA Vera Rubin Superchip

Superchip · 2026
FP32: 260.0 TFLOPS
Compare vs Vera Rubin

NVIDIA Rubin SXM

SXM · 2026
FP32: 130.0 TFLOPS
Compare vs Rubin

Stay Updated on GPU Releases

Get notified when new GPUs are added or specifications are updated.

Loading verification...

No spam, unsubscribe anytime.

Back to GPUs