Milestones
- Rubin CPX announcedOct 2026Complete.
- Vera Rubin NVL144 CPX system productionTarget Q4 2026Not yet reached.
Most important, last 3 months
Upcoming
- Q4 2026Vera Rubin NVL144 CPX system production (next)
by NVIDIAUS
GPU for million-token context windows; Vera Rubin NVL144 CPX system end of 2026
Updated Oct 2026Checked 10 Oct0 updates this week
GDDR7 is lower cost and faster to produce at scale. Suitable for prefill where throughput matters more than per-access latency.
Custom attention algorithms handle million-token contexts. Achieves 3x speedup on attention over prior generation.
NVIDIA's 4-bit float format reduces memory bandwidth and power. Suitable for prefill where numerical precision is less critical than throughput.
| Spec | Rubin CPX |
|---|---|
| NVFP4 compute per GPU | 30 petaflopsR (reported) |
| Memory per GPU | 128 GB GDDR7R (reported) |
| Attention speedup vs prior generation | 3x fasterR (reported) |
| Vera Rubin NVL144 CPX per-rack | 8 exaflopsR (reported) |
| NVL144 CPX memory bandwidth | 1.7 PB/sR (reported) |
R reported by the company
NVIDIA
AI GPUs, rack systems, CUDA, networking
NVIDIA designs the GPUs, networking and software behind most AI training and inference. Its CUDA software made GPUs the default for AI. In October 2025 it became the first company worth $5 trillion.