5 papers
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
Can Xiao, Sukmin Cho, Junbong We +7
Large language model (LLM) inference serving is increasingly constrained by memory rather than compute. As long-context and long-form reasoning workloads become more prevalent, the…
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
Haoran Wu, Zeyu Cao, Yao Lai +15
Emerging agentic LLM workloads are driving rapidly growing demand on both memory capacity and bandwidth, with different phases of inference (e.g., prefill and decode) imposing dist…
A Scalable Open-Source QEC System with Sub-Microsecond Decoding-Feedback Latency
Junyi Liu, Yi Lee, Yilun Xu +2
Quantum error correction (QEC) is essential for realizing large-scale, fault-tolerant quantum computation, yet its practical implementation remains a major engineering challenge. I…
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
Burcu Canakci, Junyi Liu, Xingbo Wu +5
To match the blooming demand of generative AI workloads, GPU designers have so far been trying to pack more and more compute and memory into single complex and expensive packages.…
Managed-Retention Memory: A New Class of Memory for the AI Era
Sergey Legtchenko, Ioan Stefanovici, Richard Black +6
AI clusters today are one of the major uses of High Bandwidth Memory (HBM). However, HBM is suboptimal for AI workloads for several reasons. Analysis shows HBM is overprovisioned o…