NVIDIA's Rubin CPX GPU Returns With HBM4 Memory

Quick Report

NVIDIA has reportedly revived its Rubin CPX AI accelerator, this time with HBM4 memory instead of the originally planned GDDR7. The redesign suggests the company is treating the platform as a specialized AI prefill engine for large-context workloads rather than a broad consumer-style accelerator.

The report says the chip will carry around 168 GB of HBM4 memory and will be deployed in racks with multiple CPX GPUs handling prefill and KV-cache operations, while regular Rubin GPUs handle decode tasks. That split is timely because long-context AI inference is increasingly memory-bandwidth constrained, and HBM4 offers the bandwidth profile needed for that kind of workload.

Written using GitHub Copilot GPT-5 mini in agentic mode instructed to follow current codebase style and conventions for writing articles.

Source(s)

  • TPU