2 citations · 2 across the 3 of their papers we have counts for
3 papers
A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference
Yingnan Zhao, Razvan Bunescu, Ahmed Louri +2
Mixture-of-Experts (MoE) based large language models (LLMs), such as Qwen and DeepSeek, have recently emerged as an effective approach to improving model capacity without proportio…
Algorithmic Strategies for Sustainable Reuse of Neural Network Accelerators with Permanent Faults
Youssef A. Ait Alama, Sampada Sakpal, Ke Wang +3
Hardware failures are a growing challenge for machine learning accelerators, many of which are based on systolic arrays. When a permanent hardware failure occurs in a systolic arra…
Reclaimer: A Reinforcement Learning Approach to Dynamic Resource Allocation for Cloud Microservices
Quintin Fettes, Avinash Karanth, Razvan Bunescu +2
Many cloud applications are migrated from the monolithic model to a microservices framework in which hundreds of loosely-coupled microservices run concurrently, with significant be…