StoryHardware

NVIDIA announces Rubin CPX, a GPU for million-token-context inference, due at the end of 2026

NVIDIAVera Rubin

NVIDIA announced Rubin CPX, a GPU purpose-built for massive-context processing: AI coding assistants analyzing entire software projects and generative video processing handling up to 1 million tokens per hour of content. It features 128GB GDDR7 memory, 30 petaflops of NVFP4 compute, and integrated video encoders/decoders, with availability expected at end of 2026.

  • Rubin CPX handles context-processing phase of inference for coding and video generation.
  • 128GB GDDR7 memory and 30 petaflops NVFP4 compute per GPU.
  • Vera Rubin NVL144 CPX rack delivers 8 exaflops compute, 100TB fast memory.
  • 3× faster attention capabilities vs. previous platforms; expected availability end of 2026.
Read the original · Press

More on Vera Rubin

Primer