Showing 2025Show all
2 papers · 1 filter
cs.DC2025
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
Yanpeng Yu, Haiyue Ma, Krish Agarwal +10
Expert Parallelism (EP) permits Mixture of Experts (MoE) models to scale beyond a single GPU. To address load imbalance across GPUs in EP, existing approaches aim to balance the nu…
cs.RO2025
L3M+P: Lifelong Planning with Large Language Models
Krish Agarwal, Yuqian Jiang, Jiaheng Hu +2
By combining classical planning methods with large language models (LLMs), recent research such as LLM+P has enabled agents to plan for general tasks given in natural language. How…