StoryHardware
NVIDIA announces Rubin CPX, a GPU for million-token-context inference, due at the end of 2026
NVIDIA announced Rubin CPX, a GPU purpose-built for massive-context processing: AI coding assistants analyzing entire software projects and generative video processing handling up to 1 million tokens per hour of content. It features 128GB GDDR7 memory, 30 petaflops of NVFP4 compute, and integrated video encoders/decoders, with availability expected at end of 2026.
- Rubin CPX handles context-processing phase of inference for coding and video generation.
- 128GB GDDR7 memory and 30 petaflops NVFP4 compute per GPU.
- Vera Rubin NVL144 CPX rack delivers 8 exaflops compute, 100TB fast memory.
- 3× faster attention capabilities vs. previous platforms; expected availability end of 2026.