1 citations · 2 across the 20 of their papers we have counts for
12 papers · 1 filter
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
Shiqi Liu, Zeyu He, Letian Tao +9
On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and op…
D-MOPD: Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
Zechen Sun, Zhiwei Zhang, Fei Zhao +7
Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollou…
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
Guojian Zhan, Letian Tao, Pengcheng Wang +6
Learning expressive and efficient policy functions is a promising direction in reinforcement learning (RL). While flow-based policies have recently proven effective in modeling com…
DADP: Domain Adaptive Diffusion Policy
Pengcheng Wang, Qinghang Liu, Haotian Lin +4
Learning domain adaptive policies that can generalize to unseen transition dynamics, remains a fundamental challenge in learning-based control. Substantial progress has been made t…
Bootstrap Off-policy with World Model
Guojian Zhan, Likun Wang, Xiangteng Zhang +3
Online planning has proven effective in reinforcement learning (RL) for improving sample efficiency and final performance. However, using planning for environment interaction inevi…
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
Guojian Zhan, Likun Wang, Pengcheng Wang +4
Maximum entropy has become a mainstream off-policy reinforcement learning (RL) framework for balancing exploitation and exploration. However, two bottlenecks still limit further pe…