2 papers
cs.LG2026
Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space
Yuchen Li, Zongqi Fan, Nguyen H. Tran +1
Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was…
cs.LG2026
Routing in Gradient Space: Balanced Usage Is Not Expert Specialization
Yuchen Li, Mingyu Du, Zongqi Fan +2
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem a…