54 citations · 55 across the 3 of their papers we have counts for
3 papers
cs.CL2023
Adaptive Gating in Mixture-of-Experts based Language Models
Jiamin Li, Qiang Su, Yitao Yang +3
Large language models, such as OpenAI's ChatGPT, have demonstrated exceptional language understanding capabilities in various NLP tasks. Sparsely activated mixture-of-experts (MoE)…
cs.DC2022★ 1 cited
Accelerating Distributed MoE Training and Inference with Lina
Jiamin Li, Yimin Jiang, Yibo Zhu +2
Scaling model parameters improves model quality at the price of high computation overhead. Sparsely activated models, usually in the form of Mixture of Experts (MoE) architecture,…
cs.DC2022★ 54 cited
Aryl: An Elastic Cluster Scheduler for Deep Learning
Jiamin Li, Hong Xu, Yibo Zhu +3
Companies build separate training and inference GPU clusters for deep learning, and use separate schedulers to manage them. This leads to problems for both training and inference:…