most citedLHNN: Lattice Hypergraph Neural Network for VLSI Congestion Prediction

8 citations · 21 across the 10 of their papers we have counts for

collaborators

11 papers

cs.LG2022

State-Aware Proximal Pessimistic Algorithms for Offline Reinforcement Learning

Chen Chen, Hongyao Tang, Yi Ma +4

Pessimism is of great importance in offline reinforcement learning (RL). One broad category of offline RL algorithms fulfills pessimism by explicit or implicit behavior regularizat…

cs.LG2022

Prototypical context-aware dynamics generalization for high-dimensional model-based reinforcement learning

Junjie Wang, Yao Mu, Dong Li +6

The latent world model provides a promising way to learn policies in a compact latent space for tasks with high-dimensional observations, however, its generalization across diverse…

cs.LG20221 cited

Decomposed Mutual Information Optimization for Generalized Context in Meta-Reinforcement Learning

Yao Mu, Yuzheng Zhuang, Fei Ni +4

Adapting to the changes in transition dynamics is essential in robotic applications. By learning a conditional policy with a compact context, context-aware meta-reinforcement learn…

cs.AI2022

On the Convergence Theory of Meta Reinforcement Learning with Personalized Policies

Haozhi Wang, Qing Wang, Yunfeng Shao +3

Modern meta-reinforcement learning (Meta-RL) methods are mainly developed based on model-agnostic meta-learning, which performs policy gradient steps across tasks to maximize polic…

cs.LG2022

Towards A Unified Policy Abstraction Theory and Representation Learning Approach in Markov Decision Processes

Min Zhang, Hongyao Tang, Jianye Hao +1

Lying on the heart of intelligent decision-making systems, how policy is represented and optimized is a fundamental problem. The root challenge in this problem is the large scale a…

cs.LG2022

PAnDR: Fast Adaptation to New Environments from Offline Experiences via Decoupling Policy and Environment Representations

Tong Sang, Hongyao Tang, Yi Ma +5

Deep Reinforcement Learning (DRL) has been a promising solution to many complex decision-making problems. Nevertheless, the notorious weakness in generalization among environments…