2 papers
cs.DC2026
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
Lu Zhao, Rong Shi, Shaoqing Zhang +21
The training of large-scale Mixture of Experts (MoE) models faces a critical memory bottleneck due to severe load imbalance caused by dynamic token routing. This imbalance leads to…
cs.CL2024
MELA: Multilingual Evaluation of Linguistic Acceptability
Ziyin Zhang, Yikang Liu, Weifang Huang +3
In this work, we present the largest benchmark to date on linguistic acceptability: Multilingual Evaluation of Linguistic Acceptability -- MELA, with 46K samples covering 10 langua…