20 citations · 67 across the 15 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
Yongqi Huang, Peng Ye, Chenyu Huang +5
Upcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into M…
cs.LG2023★ 1 cited
Stimulative Training++: Go Beyond The Performance Limits of Residual Networks
Peng Ye, Tong He, Shengji Tang +4
Residual networks have shown great success and become indispensable in recent deep neural network models. In this work, we aim to re-investigate the training process of residual ne…
cs.LG2022★ 15 cited
-DARTS: Beta-Decay Regularization for Differentiable Architecture Search
Peng Ye, Baopu Li, Yikang Li +3
Neural Architecture Search~(NAS) has attracted increasingly more attention in recent years because of its capability to design deep neural networks automatically. Among them, diffe…