3 papers
cs.LG2026
Training-Free Looped Transformers
Lizhang Chen, Jonathan Li, Chen Liang +2
We introduce training-free looped transformers, in which a lightweight inference-time wrapper loops a contiguous mid-stack block of layers of a frozen checkpoint without additional…
cs.LG2026
-Balancing for Mixture-of-Experts Training
Lizhang Chen, Jonathan Li, Qi Wang +5
Mixture-of-Experts (MoE) models rely on balanced expert utilization to fully realize their scalability. However, existing load-balancing methods are largely heuristic and operate o…
cs.LG2026
Cautious Weight Decay
Lizhang Chen, Jonathan Li, Kaizhao Liang +6
We introduce Cautious Weight Decay (CWD), a one-line, optimizer-agnostic modification that applies weight decay only to parameter coordinates whose signs align with the optimizer u…