Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Escaping the KL Agreement Trap in On-Policy Distillation
Haoran Xin, Anhao Zhao, Ying Sun +3
On-policy distillation (OPD) provides dense token-level supervision by asking a teacher to score student-generated rollouts. However, when the student drifts into an unrecoverable…
cs.LG2026
Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies
Jin Li, Yue Wu, Mengsha Huang +3
Recurrent neural policies are widely used in partially observable control and meta-RL tasks. Their abilities to maintain internal memory and adapt quickly to unseen scenarios have…
cs.LG2026
Signal-Adaptive Trust Regions for Gradient-Free Optimization of Recurrent Spiking Neural Networks
Jinhao Li, Yuhao Sun, Zhiyuan Ma +5
Recurrent spiking neural networks (RSNNs) are a promising substrate for energy-efficient control policies, but training them for high-dimensional, long-horizon reinforcement learni…