3 citations · 3 across the 1 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2024
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
Jinze Zhao, Peihao Wang, Junjie Yang +6
Sparse Mixture-of-Experts (SMoE) architectures have gained prominence for their ability to scale neural networks, particularly transformers, without a proportional increase in comp…
cs.LG2024★ 3 cited
Take the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization
Xuxi Chen, Zhendong Wang, Daouda Sow +5
In the rapidly advancing arena of large language models (LLMs), a key challenge is to enhance their capabilities amid a looming shortage of high-quality training data. Our study st…
cs.LG2023
Rethinking PGD Attack: Is Sign Function Necessary?
Junjie Yang, Tianlong Chen, Xuxi Chen +2
Neural networks have demonstrated success in various domains, yet their performance can be significantly degraded by even a small input perturbation. Consequently, the construction…