New flash-based memory could boost GPU capacity to terabytes
SanDisk and SK Hynix are developing high-bandwidth flash (HBF), a NAND-based memory that aims to deliver terabyte-scale capacities at speeds comparable to HBM for AI accelerators.
SanDisk and SK Hynix are collaborating on high-bandwidth flash (HBF), a storage-inspired memory that stacks 16 NAND layers to provide up to 512 GB per module and read speeds initially around 1.6 TB/s, with future versions targeting over 3 TB/s. Unlike traditional HBM, which uses DRAM, HBF relies on NAND flash, giving it far greater capacity but also lower write endurance and microsecond-scale latency. The firms propose a hybrid approach where HBM handles write-intensive pre-fill operations and HBF stores model weights for the read-dominant decode stage of AI inference, effectively acting as a non-volatile, read-many memory.
Sample chips are planned for release later this year, and inference hardware using HBF could appear early next year, though adoption hinges on industry standards and co-packaging with GPU or ASIC manufacturers. The technology could enable single-accelerator deployment of multi-trillion-parameter models and reduce reliance on high-speed interconnects, but practical rollout may take several years.
Why it matters
If successful, HBF could let AI systems run far larger models on fewer GPUs, lowering costs and simplifying data-center designs.
In this story