NVIDIA Rubin CPX NEW
Overview
NVIDIA Rubin CPX is a context-phase inference accelerator announced on 9 September 2025 at the AI Infrastructure Summit. It is a distinct class of part from the Rubin datacenter GPU: same Rubin architecture, but a cost-efficient monolithic die paired with 128 GB of GDDR7 instead of HBM4, on the reasoning that the prefill or context phase of long-context inference is compute bound rather than memory-bandwidth bound. NVIDIA published only three hardware figures for it: 30 PFLOPS of NVFP4 compute, 128 GB of GDDR7, and 3x the attention acceleration of GB300 NVL72, alongside hardware video encode and decode. Memory bandwidth, TDP, process node, transistor count, form factor, NVLink and PCIe generation were never disclosed and are stored as NULL here rather than estimated. CPX was designed to sit alongside Rubin GPUs in the Vera Rubin NVL144 CPX rack (144 Rubin CPX, 144 Rubin GPUs, 36 Vera CPUs, 8 exaflops NVFP4, 100 TB of memory, 1.7 PB/s of bandwidth), or in a dedicated CPX compute tray for existing Vera Rubin NVL144 installations. Both the chip and that rack disappeared from NVIDIA's roadmap at GTC 2026.
Performance Metrics
Peak theoretical throughput by precision type
| Precision | Bits | Peak TFLOPS | |
|---|---|---|---|
| NVFP4 | -- | 0.0 |
Power Specifications
TDP
--
Max Power
--
Power Connector
PCIe Slot
Cooling
Air
Memory Specifications
Capacity
128 GB
Type
GDDR7
Bandwidth
--
Interface
--
Hardware & Design
Form Factor
--
Architecture
Rubin
Process Node
--
Launch Year
2026
Variant
Standard
Market Segment
Professional
Chiplets
1 (Monolithic)
Full Specifications
| Chip Design | |
|---|---|
| Chiplets | 1 (Monolithic) |
| Memory | |
| VRAM | 128 GB |
| Memory Type | GDDR7 |
| Power & Thermal | |
| Enterprise Features | |
| Sparsity (2:4) | Yes |
| General | |
| Architecture | Rubin |
| Launch Year | 2026 |
Documentation & Resources
Common Use Cases
The NVIDIA Rubin CPX is optimized for high-performance computing tasks with Rubin architecture delivering high TFLOPS of compute power.
Where to Rent
Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.
Browse GPU Cloud ProvidersSimilar GPUs
Stay Updated on GPU Releases
Get notified when new GPUs are added or specifications are updated.
No spam, unsubscribe anytime.