2 papers
cs.AR2026
A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference
Yingnan Zhao, Razvan Bunescu, Ahmed Louri +2
Mixture-of-Experts (MoE) based large language models (LLMs), such as Qwen and DeepSeek, have recently emerged as an effective approach to improving model capacity without proportio…
cs.LG2024
Algorithmic Strategies for Sustainable Reuse of Neural Network Accelerators with Permanent Faults
Youssef A. Ait Alama, Sampada Sakpal, Ke Wang +3
Hardware failures are a growing challenge for machine learning accelerators, many of which are based on systolic arrays. When a permanent hardware failure occurs in a systolic arra…