1 paper · 1 filter
Yicheng Feng, Xin Tan, Kin Hang Sew +3
Large Language Model (LLM) inference is growing increasingly complex with the rise of Mixture-of-Experts (MoE) models and disaggregated architectures that decouple components like…