activity
20192025
most citedLearning to Utilize Shaping Rewards: A New Approach of Reward Shaping

94 citations · 155 across the 20 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2024

Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards

Zhaohui Jiang, Xuening Feng, Paul Weng +6

In practice, reinforcement learning (RL) agents are often trained with a possibly imperfect proxy reward function, which may lead to a human-agent alignment issue (i.e., the learne…

cs.LG2024

Bayesian Design Principles for Offline-to-Online Reinforcement Learning

Hao Hu, Yiqin Yang, Jianing Ye +7

Offline reinforcement learning (RL) is crucial for real-world applications where exploration can be costly or unsafe. However, offline learned policies are often suboptimal, and fu…

cs.LG2024

vMFER: Von Mises-Fisher Experience Resampling Based on Uncertainty of Gradient Directions for Policy Improvement

Yiwen Zhu, Jinyi Liu, Wenya Wei +7

Reinforcement Learning (RL) is a widely employed technique in decision-making problems, encompassing two fundamental operations -- policy evaluation and policy improvement. Enhanci…

cs.LG2023★ 1 cited

Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning

Jinyi Liu, Yi Ma, Jianye Hao +4

In recent years, data-driven reinforcement learning (RL), also known as offline RL, have gained significant attention. However, the role of data sampling techniques in offline RL h…

cs.LG2023★ 5 cited

Neural Episodic Control with State Abstraction

Zhuo Li, Derui Zhu, Yujing Hu +6

Existing Deep Reinforcement Learning (DRL) algorithms suffer from sample inefficiency. Generally, episodic control-based approaches are solutions that leverage highly-rewarded past…

cs.LG2022★ 1 cited

EUCLID: Towards Efficient Unsupervised Reinforcement Learning with Multi-choice Dynamics Model

Yifu Yuan, Jianye Hao, Fei Ni +6

Unsupervised reinforcement learning (URL) poses a promising paradigm to learn useful behaviors in a task-agnostic environment without the guidance of extrinsic rewards to facilitat…