Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
SR-OPSD: Self-Referenced On-Policy Self-Distillation
Zhuo Sun, Entong Li, Yanlong Zhao +7
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to re…
cs.LG2026
Fisher Decorator: Refining Flow Policy via a Local Transport Map
Xiaoyuan Cheng, Haoyu Wang, Wenxuan Yuan +4
Recent advances in flow-based offline reinforcement learning (RL) have achieved strong performance by parameterizing policies via flow matching. However, they still face critical t…