2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.DC2025★ 2 cited
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
Yongji Wu, Xueshen Liu, Shuowei Jin +6
The Mixture-of-Experts (MoE) architecture has become increasingly popular as a method to scale up large language models (LLMs). To save costs, heterogeneity-aware training solution…
cs.CL2024
Plato: Plan to Efficiently Decode for Large Language Model Inference
Shuowei Jin, Xueshen Liu, Yongji Wu +7
Large language models (LLMs) have achieved remarkable success in natural language tasks, but their inference incurs substantial computational and memory overhead. To improve effici…