4 papers
SR-OPSD: Self-Referenced On-Policy Self-Distillation
Zhuo Sun, Entong Li, Yanlong Zhao +7
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to re…
Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation
Yaowei Zheng, Richong Zhang, Shenxi Wu +5
We study finite-horizon continuous-time policy evaluation from discrete closed-loop trajectories under time-inhomogeneous dynamics. The target value surface solves a backward parab…
Fisher Decorator: Refining Flow Policy via a Local Transport Map
Xiaoyuan Cheng, Haoyu Wang, Wenxuan Yuan +4
Recent advances in flow-based offline reinforcement learning (RL) have achieved strong performance by parameterizing policies via flow matching. However, they still face critical t…
CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models
Ryan Feng Lin, Keyu Tian, Hanming Zheng +3
The emergence of large language models (LLMs) has sparked much interest in creating LLM-based digital populations that can be applied to many applications such as social simulation…