5 citations · 14 across the 8 of their papers we have counts for
4 papers · 1 filter
Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding
Jiahao Zheng, Yifan Qin, Xiaobo Sharon Hu +1
The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so decoding is memory-bandwidth bound.…
A 10.60 W 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia Detection
Yifan Qin, Zhenge Jia, Zheyu Yan +9
This paper proposes an ultra-low power, mixed-bit-width sparse convolutional neural network (CNN) accelerator to accelerate ventricular arrhythmia (VA) detection. The chip achieves…
TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM Accelerators
Yifan Qin, Zheyu Yan, Zixuan Pan +3
Compute-in-memory (CIM) accelerators using non-volatile memory (NVM) devices offer promising solutions for energy-efficient and low-latency Deep Neural Network (DNN) inference exec…
A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity
Yuxiang Huan, Yifan Qin, Yantian You +2
It remains a challenge to run Deep Learning in devices with stringent power budget in the Internet-of-Things. This paper presents a low-power accelerator for processing Deep Neural…