2 papers
cs.AR2026
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
Ye Lin, Chao Fang, Xiaoyong Song +4
Edge deployment of low-batch large language models (LLMs) faces critical memory bandwidth bottlenecks when executing memory-intensive general matrix-vector multiplications (GEMV) o…
cs.DC2026
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
Qi Wu, Chao Fang, Jiayuan Chen +5
Mixture-of-Experts (MoE) models facilitate edge deployment by decoupling model capacity from active computation, yet their large memory footprint drives the need for GPU systems wi…