Semiconductor Engineering reports that SK hynix researchers published a July 2026 technical paper on StreamDQ, a custom high-bandwidth memory concept for LLM inference. The work targets weight dequantization inside the memory subsystem rather than leaving that step entirely to the GPU compute path. The paper reports large gains for mixed-precision GEMM workloads, including up to 7.08x speedup and 90.23% lower energy in the evaluated cases. For RamTrend, the market signal is that HBM vendors are still looking beyond raw bandwidth and capacity toward architecture-level differentiation for AI inference. This is research coverage rather than a product launch, so it should not be read as evidence of near-term supply or pricing changes.
AI Memory · Jul 13, 2026
SK Hynix Paper Points To Custom HBM For Faster LLM Inference
SK hynix researchers have published a paper on StreamDQ, a custom-HBM architecture that moves weight dequantization closer to memory for large-language-model inference workloads.
Price impact: 1Direction: upSource: Semiconductor Engineering
SK hynixHBMcustom HBMnear-memory computingLLM inferenceAI memory
Original sourceBack to news archive