1 paper
Hyunsun Chung, Taewan Noh, Minji Kim +3
NAND-backed storage offers the capacity needed to scale LLM prefix caching, but its block I/O path incurs CPU cache contention and host-DRAM staging in addition to NAND latency. Ou…