4 papers
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
Cenlin Duan, Jianlei Yang, Rubing Yang +8
The deployment of large language models (LLMs) presents significant challenges due to their enormous memory footprints, low arithmetic intensity, and stringent latency requirements…
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
Tong Qiao, Ao Zhou, Yingjie Qi +4
Graph Neural Networks (GNNs) have been widely adopted due to their strong performance. However, GNN training often relies on expensive, high-performance computing platforms, limiti…
Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
Cenlin Duan, Jianlei Yang, Yikun Wang +7
Processing-in-memory (PIM) is a transformative architectural paradigm designed to overcome the Von Neumann bottleneck. Among PIM architectures, digital SRAM-PIM emerges as a promis…
CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM Architectures
Yingjie Qi, Jianlei Yang, Yiou Wang +6
Digital Compute-in-Memory (CIM) architectures have shown great promise in Deep Neural Network (DNN) acceleration by effectively addressing the "memory wall" bottleneck. However, th…