activity
20192022
most citedRethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL

13 citations · 29 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG20222 cited

Optimistic Curiosity Exploration and Conservative Exploitation with Linear Reward Shaping

Hao Sun, Lei Han, Rui Yang +3

In this work, we study the simple yet universally applicable case of reward shaping in value-based Deep Reinforcement Learning (DRL). We show that reward shifting in the form of th…

cs.LG202213 cited

Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL

Rui Yang, Yiming Lu, Wenzhe Li +6

Solving goal-conditioned tasks with sparse rewards using self-supervised learning is promising because of its simplicity and stability over current reinforcement learning (RL) algo…

cs.LG20212 cited

Safe Exploration by Solving Early Terminated MDP

Hao Sun, Ziping Xu, Meng Fang +4

Safe exploration is crucial for the real-world application of reinforcement learning (RL). Previous works consider the safe exploration problem as Constrained Markov Decision Proce…

cs.LG20205 cited

Non-local Policy Optimization via Diversity-regularized Collaborative Exploration

Zhenghao Peng, Hao Sun, Bolei Zhou

Conventional Reinforcement Learning (RL) algorithms usually have one single agent learning to solve the task independently. As a result, the agent can only explore a limited part o…

cs.LG2020

Evolutionary Stochastic Policy Distillation

Hao Sun, Xinyu Pan, Bo Dai +2

Solving the Goal-Conditioned Reward Sparse (GCRS) task is a challenging reinforcement learning problem due to the sparsity of reward signals. In this work, we propose a new formula…

cs.LG2019

Risk-Averse Trust Region Optimization for Reward-Volatility Reduction

Lorenzo Bisi, Luca Sabbioni, Edoardo Vittori +2

In real-world decision-making problems, for instance in the fields of finance, robotics or autonomous driving, keeping uncertainty under control is as important as maximizing expec…