2 papers
cs.AR2026
DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference
Xiaoyang Lu, Belthangady Akash Vi Narayana Pai, Xian-He Sun
Mixture-of-Experts (MoE) models enable efficient scaling of large language model (LLM) inference but suffer from substantial data-movement overhead when deployed on neural processi…
cs.AR2026
VIPER: Architecture-Aware Performance Modeling for Processing-in-Memory Design-Space Exploration
Haoran Geng, Tomas Sousa Pereira, Xiaoyang Lu +3
Processing-in-Memory (PIM) promises to reduce data movement overhead by executing computation in or near memory, but its realized application speedup remains highly design-dependen…