3 papers
cs.AI2026
Off-Policy Actor-Critic with Sigmoid-Bounded Entropy for Real-World Robot Learning
Xiefeng Wu, Mingyu Hu, Shu Zhang
Deploying reinforcement learning in the real world remains challenging due to sample inefficiency, sparse rewards, and noisy visual observations. Prior work leverages demonstration…
cs.LG2025
Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning
Xiefeng Wu, Jing Zhao, Shu Zhang +1
Online reinforcement learning in complex tasks is time-consuming, as massive interaction steps are needed to learn the optimal Q-function.Vision-language action (VLA) policies repr…
cs.AI2024
From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge
Xiefeng Wu
Q-shaping is an extension of Q-value initialization and serves as an alternative to reward shaping for incorporating domain knowledge to accelerate agent training, thereby improvin…