7 citations · 12 across the 7 of their papers we have counts for
6 papers · 1 filter
FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
Yuwei Fu, Haichao Zhang, Di Wu +2
In this work, we investigate how to leverage pre-trained visual-language models (VLM) for online Reinforcement Learning (RL). In particular, we focus on sparse reward tasks with pr…
PaCo: Parameter-Compositional Multi-Task Reinforcement Learning
Lingfeng Sun, Haichao Zhang, Wei Xu +1
The purpose of multi-task reinforcement learning (MTRL) is to train a single policy that can be applied to a set of different tasks. Sharing parameters allows us to take advantage…
Generative Planning for Temporally Coordinated Exploration in Reinforcement Learning
Haichao Zhang, Wei Xu, Haonan Yu
Standard model-free reinforcement learning algorithms optimize a policy that generates the action to be taken in the current time step in order to maximize expected future return.…
Do You Need the Entropy Reward (in Practice)?
Haonan Yu, Haichao Zhang, Wei Xu
Maximum entropy (MaxEnt) RL maximizes a combination of the original task reward and an entropy reward. It is believed that the regularization imposed by entropy, on both policy imp…
Feature Importance in a Deep Learning Climate Emulator
Wei Xu, Xihaier Luo, Yihui Ren +3
We present a study using a class of post-hoc local explanation methods i.e., feature importance methods for "understanding" a deep learning (DL) emulator of climate. Specifically,…
TAAC: Temporally Abstract Actor-Critic for Continuous Control
Haonan Yu, Wei Xu, Haichao Zhang
We present temporally abstract actor-critic (TAAC), a simple but effective off-policy RL algorithm that incorporates closed-loop temporal abstraction into the actor-critic framewor…