StoryHardware

NVIDIA unveils Rubin CPX, a GPU optimized for million-token contexts

NVIDIARubin CPX

NVIDIA announced Rubin CPX, a GPU for prefill on massive context windows up to one million tokens. It delivers 30 petaflops NVFP4 compute and 128GB GDDR7 memory per GPU. The Vera Rubin NVL144 CPX rack achieves 8 exaflops and 1.7 PB/s memory bandwidth.

  • Addresses a gap: GPUs handle prefill only to ~100k tokens, but many applications need longer contexts.
  • Uses GDDR7 instead of HBM to trade bandwidth for cost and manufacturability at scale.
  • Enables AI coding assistants to ingest entire codebases and video processing of hours of content in a single context.
  • Expected availability end of 2026.
Read the original · Press
Primer