activity
20162020
most citedRobust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations

9 citations · 10 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG20201 cited

Stabilizing Q Learning Via Soft Mellowmax Operator

Yaozhong Gan, Zhe Zhang, Xiaoyang Tan

Learning complicated value functions in high dimensional state space by function approximation is a challenging task, partially due to that the max-operator used in temporal differ…

cs.CV2020

Deep Robust Multilevel Semantic Cross-Modal Hashing

Ge Song, Jun Zhao, Xiaoyang Tan

Hashing based cross-modal retrieval has recently made significant progress. But straightforward embedding data from different modalities into a joint Hamming space will inevitably…

cs.MA2019

SMIX(): Enhancing Centralized Value Functions for Cooperative Multi-Agent Reinforcement Learning

Xinghu Yao, Chao Wen, Yuhui Wang +1

Learning a stable and generalizable centralized value function (CVF) is a crucial but challenging task in multi-agent reinforcement learning (MARL), as it has to deal with the issu…

cs.LG2019

Truly Proximal Policy Optimization

Yuhui Wang, Hao He, Chao Wen +1

Proximal policy optimization (PPO) is one of the most successful deep reinforcement-learning methods, achieving state-of-the-art performance across a wide range of challenging task…

cs.LG20199 cited

Robust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations

Yuhui Wang, Hao He, Xiaoyang Tan

In real-world scenarios, the observation data for reinforcement learning with continuous control is commonly noisy and part of it may be dynamically missing over time, which violat…

cs.LG2019

Trust Region-Guided Proximal Policy Optimization

Yuhui Wang, Hao He, Xiaoyang Tan +1

Proximal policy optimization (PPO) is one of the most popular deep reinforcement learning (RL) methods, achieving state-of-the-art performance across a wide range of challenging ta…