2 papers
cs.AI2026
Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective
Junpeng Zhang, Lei Cheng, Guoxi Zhang +3
This paper explores a scientific question in supervised fine-tuning (SFT): why SFT is broadly effective for small-scale deep neural networks, yet can produce inconsistent or even d…
cs.LG2026
SPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement Learning
Lirui Luo, Guoxi Zhang, Hongming Xu +2
In deep reinforcement learning (DRL), an agent is trained from a stream of experience. In a continual learning setting, such agents can suffer from plasticity loss: their ability t…