16 papers
Diagnosing Compositional Generalization in Sequential Robot Tasks
Yixiao Wang, Cheng-En Wu, Lingfeng Sun +5
Sequential robot manipulation requires policies to execute novel combinations of familiar instruction components. However, collecting demonstrations for all possible instruction tu…
Bridging the Gap between Newton-Raphson Method and Regularized Policy Iteration
Zeyang Li, Chuxiong Hu, Yunan Wang +4
The paper shows that regularized policy iteration in reinforcement learning is mathematically equivalent to applying the Newton‑Raphson method to a smoothed Bellman equation, provi…
Factor-Aware Mixture-of-Experts with Pretrained Encoder for Combinatorial Generalization
Feihong Zhang, Guojian Zhan, Zeyu He +8
The integration of pretrained encoders with diffusion policies has become a dominant paradigm for visual robotic manipulation. However, it still struggles to generalize across comp…
DADP: Domain Adaptive Diffusion Policy
Pengcheng Wang, Qinghang Liu, Haotian Lin +4
Learning domain adaptive policies that can generalize to unseen transition dynamics, remains a fundamental challenge in learning-based control. Substantial progress has been made t…
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
Shiqi Liu, Zeyu He, Guojian Zhan +10
Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regu…
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
Guojian Zhan, Letian Tao, Pengcheng Wang +6
Learning expressive and efficient policy functions is a promising direction in reinforcement learning (RL). While flow-based policies have recently proven effective in modeling com…