collaborators

5 papers

cs.DC2026

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

Can Xiao, Sukmin Cho, Junbong We +7

Large language model (LLM) inference serving is increasingly constrained by memory rather than compute. As long-context and long-form reasoning workloads become more prevalent, the…

cs.AR2026

MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs

Haoran Wu, Zeyu Cao, Yao Lai +15

Emerging agentic LLM workloads are driving rapidly growing demand on both memory capacity and bandwidth, with different phases of inference (e.g., prefill and decode) imposing dist…

quant-ph2026

A Scalable Open-Source QEC System with Sub-Microsecond Decoding-Feedback Latency

Junyi Liu, Yi Lee, Yilun Xu +2

Quantum error correction (QEC) is essential for realizing large-scale, fault-tolerant quantum computation, yet its practical implementation remains a major engineering challenge. I…

cs.AR2025

Good things come in small packages: Should we build AI clusters with Lite-GPUs?

Burcu Canakci, Junyi Liu, Xingbo Wu +5

To match the blooming demand of generative AI workloads, GPU designers have so far been trying to pack more and more compute and memory into single complex and expensive packages.…

cs.AR2025

Managed-Retention Memory: A New Class of Memory for the AI Era

Sergey Legtchenko, Ioan Stefanovici, Richard Black +6

AI clusters today are one of the major uses of High Bandwidth Memory (HBM). However, HBM is suboptimal for AI workloads for several reasons. Analysis shows HBM is overprovisioned o…