247 citations · 379 across the 29 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Turn Waste into Worth: Rectifying Top- Router of MoE
Zhiyuan Zeng, Qipeng Guo, Zhaoye Fei +7
Sparse Mixture of Experts (MoE) models are popular for training large language models due to their computational efficiency. However, the commonly used top- routing mechanism su…
cs.LG2023
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
Kai Lv, Hang Yan, Qipeng Guo +2
Large language models have achieved remarkable success, but their extensive parameter size necessitates substantial memory for training, thereby setting a high threshold. While the…