4 papers
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
Jiawei Hao, Zhiwei Hao, Jianyuan Guo +4
Mixture-of-Experts (MoE) based Large Language Models (LLMs) have demonstrated impressive performance and computational efficiency. However, their deployment is often constrained by…
Maximizing Incremental Information Entropy for Contrastive Learning
Jiansong Zhang, Zhuoqin Yang, Xu Wu +3
Contrastive learning has achieved remarkable success in self-supervised representation learning, often guided by information-theoretic objectives such as mutual information maximiz…
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
Zhiwei Hao, Jianyuan Guo, Li Shen +6
Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources required for their training present a signific…
SeWA: Selective Weight Average via Probabilistic Masking
Peng Wang, Shengchao Hu, Zerui Tao +5
Weight averaging has become a standard technique for enhancing model performance. However, methods such as Stochastic Weight Averaging (SWA) and Latest Weight Averaging (LAWA) ofte…