collaborators

6 papers

cs.AR2026

HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference

Cenlin Duan, Jianlei Yang, Rubing Yang +8

The deployment of large language models (LLMs) presents significant challenges due to their enormous memory footprints, low arithmetic intensity, and stringent latency requirements…

cs.LG2025

GCoDE: Efficient Device-Edge Co-Inference for GNNs via Architecture-Mapping Co-Search

Ao Zhou, Jianlei Yang, Tong Qiao +4

Graph Neural Networks (GNNs) have emerged as the state-of-the-art graph learning method. However, achieving efficient GNN inference on edge devices poses significant challenges, li…

cs.LG2025

TinyFormer: Efficient Transformer Design and Deployment on Tiny Devices

Jianlei Yang, Jiacheng Liao, Fanding Lei +6

Developing deep learning models on tiny devices (e.g. Microcontroller units, MCUs) has attracted much attention in various embedded IoT applications. However, it is challenging to…

cs.AR2025

CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-based CIM Architectures

Yingjie Qi, Jianlei Yang, Rubing Yang +5

Compute-in-memory (CIM) has emerged as a pivotal direction for accelerating workloads in the field of machine learning, such as Deep Neural Networks (DNNs). However, the effective…

cs.AR2025

Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity

Cenlin Duan, Jianlei Yang, Yikun Wang +7

Processing-in-memory (PIM) is a transformative architectural paradigm designed to overcome the Von Neumann bottleneck. Among PIM architectures, digital SRAM-PIM emerges as a promis…

cs.AR2025

CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM Architectures

Yingjie Qi, Jianlei Yang, Yiou Wang +6

Digital Compute-in-Memory (CIM) architectures have shown great promise in Deep Neural Network (DNN) acceleration by effectively addressing the "memory wall" bottleneck. However, th…