4 papers · 1 filter
MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models
Chenglin Liu, Xun Wang, Ruishuo Chen +2
Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and re…
When Context Returns: Toward Robust Internalization in On-Policy Distillation
Xun Wang, Ruishuo Chen, Zhuoran Li +2
Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer ne…
Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms
Tian Xu, Zhilong Zhang, Zexuan Chen +3
Adversarial imitation learning (AIL), a prominent approach in imitation learning, has achieved significant practical success powered by neural network approximation. However, exist…
Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
Tian Xu, Zhilong Zhang, Ruishuo Chen +2
As a prominent category of imitation learning methods, adversarial imitation learning (AIL) has garnered significant practical success powered by neural network approximation. Howe…