4 papers
On the Identifiability of Controlled World Models
Xiangteng Zhang, Yang Guan, Bo Zhang +3
World model serves as a promising tool to infer environment dynamics under high-dimensional observations and candidate actions. Recently, LeCun's JEPA provides a compelling framewo…
Bootstrap Off-policy with World Model
Guojian Zhan, Likun Wang, Xiangteng Zhang +3
Online planning has proven effective in reinforcement learning (RL) for improving sample efficiency and final performance. However, using planning for environment interaction inevi…
Off-policy Reinforcement Learning with Model-based Exploration Augmentation
Likun Wang, Xiangteng Zhang, Yinuo Wang +5
Exploration is fundamental to reinforcement learning (RL), as it determines how effectively an agent discovers and exploits the underlying structure of its environment to achieve o…
Conformal Symplectic Optimization for Stable Reinforcement Learning
Yao Lyu, Xiangteng Zhang, Shengbo Eben Li +5
Training deep reinforcement learning (RL) agents necessitates overcoming the highly unstable nonconvex stochastic optimization inherent in the trial-and-error mechanism. To tackle…