26 citations · 90 across the 20 of their papers we have counts for
8 papers · 1 filter
NVCiM-PT: An NVCiM-assisted Prompt Tuning Framework for Edge LLMs
Ruiyang Qin, Pengyu Ren, Zheyu Yan +7
Large Language Models (LLMs) deployed on edge devices, known as edge LLMs, need to continuously fine-tune their model parameters from user-generated data under limited resource con…
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
Ruiyang Qin, Dancheng Liu, Gelei Xu +7
The combination of Large Language Models (LLM) and Automatic Speech Recognition (ASR), when deployed on edge devices (called edge ASR-LLM), can serve as a powerful personalized ass…
A 10.60 W 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia Detection
Yifan Qin, Zhenge Jia, Zheyu Yan +9
This paper proposes an ultra-low power, mixed-bit-width sparse convolutional neural network (CNN) accelerator to accelerate ventricular arrhythmia (VA) detection. The chip achieves…
Rethinking Medical Anomaly Detection in Brain MRI: An Image Quality Assessment Perspective
Zixuan Pan, Jun Xia, Zheyu Yan +7
Reconstruction-based methods, particularly those leveraging autoencoders, have been widely adopted for anomaly detection task in brain MRI. Unlike most existing works try to improv…
TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM Accelerators
Yifan Qin, Zheyu Yan, Zixuan Pan +3
Compute-in-memory (CIM) accelerators using non-volatile memory (NVM) devices offer promising solutions for energy-efficient and low-latency Deep Neural Network (DNN) inference exec…
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
Ruiyang Qin, Dancheng Liu, Chenhui Xu +9
The scaling laws have become the de facto guidelines for designing large language models (LLMs), but they were studied under the assumption of unlimited computing resources for bot…