3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CL2025
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
Yehui Tang, Yichun Yin, Yaoyuan Wang +71
Sparse large language models (LLMs) with Mixture of Experts (MoE) and close to a trillion parameters are dominating the realm of most capable language models. However, the massive…
cs.LG2024
Complexity Matters: Dynamics of Feature Learning in the Presence of Spurious Correlations
GuanWen Qiu, Da Kuang, Surbhi Goel
Existing research often posits spurious features as easier to learn than core features in neural network optimization, but the impact of their relative simplicity remains under-exp…
cs.CV2022★ 3 cited
Evaluating Out-of-Distribution Performance on Document Image Classifiers
Stefan Larson, Gordon Lim, Yutong Ai +2
The ability of a document classifier to handle inputs that are drawn from a distribution different from the training distribution is crucial for robust deployment and generalizabil…