2 papers
cs.LG2026
MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
Lianhai Ren, Yucheng Ding, Xiao Liu +3
Training instability remains a critical challenge in large language model (LLM) pretraining, often manifesting as sudden gradient explosions that waste significant computational re…
cs.LG2024
Unifying back-propagation and forward-forward algorithms through model predictive control
Lianhai Ren, Qianxiao Li
We introduce a Model Predictive Control (MPC) framework for training deep neural networks, systematically unifying the Back-Propagation (BP) and Forward-Forward (FF) algorithms. At…