9 citations · 10 across the 2 of their papers we have counts for
7 papers
Stabilizing Q Learning Via Soft Mellowmax Operator
Yaozhong Gan, Zhe Zhang, Xiaoyang Tan
Learning complicated value functions in high dimensional state space by function approximation is a challenging task, partially due to that the max-operator used in temporal differ…
Deep Robust Multilevel Semantic Cross-Modal Hashing
Ge Song, Jun Zhao, Xiaoyang Tan
Hashing based cross-modal retrieval has recently made significant progress. But straightforward embedding data from different modalities into a joint Hamming space will inevitably…
SMIX(): Enhancing Centralized Value Functions for Cooperative Multi-Agent Reinforcement Learning
Xinghu Yao, Chao Wen, Yuhui Wang +1
Learning a stable and generalizable centralized value function (CVF) is a crucial but challenging task in multi-agent reinforcement learning (MARL), as it has to deal with the issu…
Truly Proximal Policy Optimization
Yuhui Wang, Hao He, Chao Wen +1
Proximal policy optimization (PPO) is one of the most successful deep reinforcement-learning methods, achieving state-of-the-art performance across a wide range of challenging task…
Robust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations
Yuhui Wang, Hao He, Xiaoyang Tan
In real-world scenarios, the observation data for reinforcement learning with continuous control is commonly noisy and part of it may be dynamically missing over time, which violat…
Trust Region-Guided Proximal Policy Optimization
Yuhui Wang, Hao He, Xiaoyang Tan +1
Proximal policy optimization (PPO) is one of the most popular deep reinforcement learning (RL) methods, achieving state-of-the-art performance across a wide range of challenging ta…