1 citations · 1 across the 12 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Multi-Gate Residuals
Zhizhan Zheng, Feiyun Zhang, Shuchun Liu +4
While Attention Residuals has shown some effectiveness in addressing the widespread issue of unbounded activation growth across deep residual layers, it inevitably incurs significa…
cs.LG2026
A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs
Zijie Liu, Jie Peng, Jinhao Duan +7
Sparse Mixture-of-Experts (SMoE) architectures are increasingly used to scale large language models efficiently, delivering strong accuracy under fixed compute budgets. However, SM…
cs.LG2024
APOLLO: SGD-like Memory, AdamW-level Performance
Hanqing Zhu, Zhenyu Zhang, Wenyan Cong +7
Large language models (LLMs) are notoriously memory-intensive during training, particularly with the popular AdamW optimizer. This memory burden necessitates using more or higher-e…