2 papers
cs.DC2026
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
Lu Zhao, Rong Shi, Shaoqing Zhang +21
The training of large-scale Mixture of Experts (MoE) models faces a critical memory bottleneck due to severe load imbalance caused by dynamic token routing. This imbalance leads to…
cs.CV2025
ClearGCD: Mitigating Shortcut Learning For Robust Generalized Category Discovery
Kailin Lyu, Jianwei He, Long Xiao +4
In open-world scenarios, Generalized Category Discovery (GCD) requires identifying both known and novel categories within unlabeled data. However, existing methods often suffer fro…