27 citations · 86 across the 23 of their papers we have counts for
Showing 2021 · cs.CLShow all
3 papers · 2 filters
cs.CL2021★ 27 cited
Taming Sparsely Activated Transformer with Stochastic Experts
Simiao Zuo, Xiaodong Liu, Jian Jiao +5
Sparsely activated models (SAMs), such as Mixture-of-Experts (MoE), can easily scale to have outrageously large amounts of parameters without significant increase in computational…
cs.CL2021★ 2 cited
Self-Training with Differentiable Teacher
Simiao Zuo, Yue Yu, Chen Liang +5
Self-training achieves enormous success in various semi-supervised and weakly-supervised learning tasks. The method can be interpreted as a teacher-student framework, where the tea…
cs.CL2021
ARCH: Efficient Adversarial Regularized Training with Caching
Simiao Zuo, Chen Liang, Haoming Jiang +5
Adversarial regularization can improve model generalization in many natural language processing tasks. However, conventional approaches are computationally expensive since they nee…