2 papers
cs.LG2026
Fast MoE Inference via Predictive Prefetching and Expert Replication
Ankit Jyothish, Ali Jannesari, Aishwarya Sarkar +1
The Mixture of Experts (MoE) architecture has become a fundamental building block in state-of-the-art large language models (LLMs), improving domain-specific expertise in LLMs and…
cs.LG2025
Enhanced Soups for Graph Neural Networks
Joseph Zuber, Aishwarya Sarkar, Joseph Jennings +1
Graph Neural Networks (GNN) have demonstrated state-of-the-art performance in numerous scientific and high-performance computing (HPC) applications. Recent work suggests that "soup…