RamTrend

HBM · Jul 22, 2026

Nvidia Rubin Inference Design Highlights Rising HBM4 Density in AI Racks

Nvidia outlined inference-focused changes in its Rubin platform, including a GPU design with 288GB of HBM4 and rack-scale efficiency upgrades. For memory markets, the update reinforces how next-generation AI systems are pushing both HBM content and bandwidth requirements higher.

Price impact: 3Direction: upSource: Tom's Hardware

Nvidia disclosed additional architectural details for its upcoming Rubin platform as it positions the design for large-scale AI inference deployments. The company said a Rubin GPU package combines two compute dies and carries 288GB of HBM4 with 22 TB/s of bandwidth, while the full NVL72 rack configuration uses 72 Rubin GPUs and 36 Vera CPUs. The company emphasized several inference-oriented changes, including updates to its Tensor Memory Accelerator for mixture-of-experts models, higher Tensor Core throughput on key operations, faster softmax handling for lower-precision formats, finer-grained kernel dependency management and more efficient inter-GPU communication over NVLink. Nvidia argues these changes should improve utilization and lower inference cost per token at rack scale. For RamTrend, the most relevant point is the continued increase in premium memory content per accelerator platform. The article does not provide new supply or pricing data for DRAM or NAND, but it supports the view that advanced AI hardware roadmaps remain a structural demand driver for HBM and associated server memory subsystems.

NvidiaHBM4HBMNVLinkAI AcceleratorsServer Memory
Original sourceBack to news archive