4 papers
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
Ye Lin, Chao Fang, Xiaoyong Song +4
Edge deployment of low-batch large language models (LLMs) faces critical memory bandwidth bottlenecks when executing memory-intensive general matrix-vector multiplications (GEMV) o…
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
Qi Wu, Chao Fang, Jiayuan Chen +5
Mixture-of-Experts (MoE) models facilitate edge deployment by decoupling model capacity from active computation, yet their large memory footprint drives the need for GPU systems wi…
A Robust Framework for Graph-based Two-Sample Tests Using Weights
Yichuan Bai, Lynna Chu
Graph-based tests are a class of non-parametric two-sample tests useful for analyzing high-dimensional data. The test statistics are constructed from similarity graphs (such as K-m…
FBQuant: FeedBack Quantization for Large Language Models
Yijiang Liu, Hengyu Fang, Liulu He +4
Deploying Large Language Models (LLMs) on edge devices is increasingly important, as it eliminates reliance on network connections, reduces expensive API calls, and enhances user p…