4 papers
HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents
Jiangze Yan, Yi Shen, Wenjing Zhang +5
Long-horizon agents rely on memory mechanisms to compress interaction history, but optimizing memory writing faces a distinct credit assignment challenge: a memory update may be re…
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
Jialei Chen, Kai Wang, Kang Chen +9
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change th…
One-Step Flow Policy Mirror Descent
Tianyi Chen, Haitong Ma, Na Li +2
Diffusion policies have achieved great success in online reinforcement learning (RL) due to their strong expressive capacity. However, the inference of diffusion policy models reli…
Efficient Online Reinforcement Learning for Diffusion Policy
Haitong Ma, Tianyi Chen, Kai Wang +2
Diffusion policies have achieved superior performance in imitation learning and offline reinforcement learning (RL) due to their rich expressiveness. However, the conventional diff…