activity
20172022
most citedEfficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement Learning

8 citations · 21 across the 5 of their papers we have counts for

collaborators

7 papers

cs.LG20226 cited

Adversarial Auto-Augment with Label Preservation: A Representation Learning Principle Guided Approach

Kaiwen Yang, Yanchao Sun, Jiahao Su +5

Data augmentation is a critical contributing factor to the success of deep learning but heavily relies on prior domain knowledge which is not always available. Recent works on auto…

cs.LG20222 cited

Distributional Reward Estimation for Effective Multi-Agent Deep Reinforcement Learning

Jifeng Hu, Yanchao Sun, Hechang Chen +4

Multi-agent reinforcement learning has drawn increasing attention in practice, e.g., robotics and automatic driving, as it can explore optimal policies using samples generated by i…

cs.LG20228 cited

Efficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement Learning

Yongyuan Liang, Yanchao Sun, Ruijie Zheng +1

Recent studies reveal that a well-trained deep reinforcement learning (RL) policy can be particularly vulnerable to adversarial perturbations on input observations. Therefore, it i…

cs.LG2020

TempLe: Learning Template of Transitions for Sample Efficient Multi-task RL

Yanchao Sun, Xiangyu Yin, Furong Huang

Transferring knowledge among various environments is important to efficiently learn multiple tasks online. Most existing methods directly use the previously learned models or previ…

cs.LG2020

Understanding Generalization in Deep Learning via Tensor Methods

Jingling Li, Yanchao Sun, Jiahao Su +2

Deep neural networks generalize well on unseen data though the number of parameters often far exceeds the number of training examples. Recently proposed complexity measures have pr…

cs.LG20191 cited

Can Agents Learn by Analogy? An Inferable Model for PAC Reinforcement Learning

Yanchao Sun, Furong Huang

Model-based reinforcement learning algorithms make decisions by building and utilizing a model of the environment. However, none of the existing algorithms attempts to infer the dy…