VAST Data said its broader collaboration with AMD now covers a reference stack for AI training and inference built around 6th Gen EPYC CPUs, Instinct GPUs, ROCm, networking hardware, and the VAST AI Operating System. The companies are positioning the platform for persistent inference and agent-style workloads that need large context windows, heavy data movement, and higher concurrency than traditional training clusters. The most relevant memory-market angle is KV cache offload. Instead of holding all inference context inside limited GPU memory, the design uses shared high-performance storage so active workloads can keep more GPU capacity available while still retaining prior context. VAST also said early testing with an Instinct MI355X system showed faster time to first token and higher token throughput when cache data was offloaded, although those results depend on system configuration and workload conditions. The hardware roadmap adds PCIe Gen6 through AMD EPYC 9006 support and uses NVMe SSD-based storage clusters plus AMD Pensando Pollara 400 networking to move data between GPUs and shared storage. If this model gains adoption, it would strengthen demand for low-latency enterprise SSD capacity and storage fabrics in AI deployments, even if it does not directly change DRAM pricing in the near term.
AI Infrastructure · Jul 28, 2026
VAST and AMD Push KV Cache Offload Toward NVMe-Centric AI Inference
VAST Data and AMD are expanding their AI infrastructure work around external KV cache handling, pairing Instinct GPUs with EPYC processors, shared storage, and high-speed networking. For memory and storage markets, the announcement matters because it shifts some long-context inference pressure away from on-package memory and toward fast SSD-backed data paths.
Price impact: 3Direction: upSource: StorageReview
VAST DataAMDDriveNetsKV cacheNVMe SSDPCIe Gen6AMD InstinctAMD EPYCROCm
Original sourceBack to news archive