3 papers
cs.AR2026
Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding
Jiahao Zheng, Yifan Qin, Xiaobo Sharon Hu +1
The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so decoding is memory-bandwidth bound.…
cs.AR2026
Probabilistic Memory for Trustworthy Edge Intelligence
Likai Pei, Jiahao Zheng, Xueji Zhao +9
Probabilistic computation plays an important role in trustworthy edge intelligence to quantify uncertainty, enhance robustness, reconstruct data, and protect privacy, but its adopt…
cs.LG2026
When Small Variations Become Big Failures: Reliability Challenges in Compute-in-Memory Neural Accelerators
Yifan Qin, Jiahao Zheng, Zheyu Yan +3
Compute-in-memory (CiM) architectures promise significant improvements in energy efficiency and throughput for deep neural network acceleration by alleviating the von Neumann bottl…