3 citations · 6 across the 3 of their papers we have counts for
4 papers
POPO: Pessimistic Offline Policy Optimization
Qiang He, Xinwen Hou
Offline reinforcement learning (RL), also known as batch RL, aims to optimize policy from a large pre-recorded dataset without interaction with the environment. This setting offers…
Highway Transformer: Self-Gating Enhanced Self-Attentive Networks
Yekun Chai, Shuo Jin, Xinwen Hou
Self-attention mechanisms have made striking state-of-the-art (SOTA) progress in various sequence learning tasks, standing on the multi-headed dot product attention by attending to…
Inducing Cooperation via Team Regret Minimization based Multi-Agent Deep Reinforcement Learning
Runsheng Yu, Zhenyu Shi, Xinrun Wang +5
Existing value-factorized based Multi-Agent deep Reinforce-ment Learning (MARL) approaches are well-performing invarious multi-agent cooperative environment under thecen-tralized t…
Learning Representations in Reinforcement Learning:An Information Bottleneck Approach
Pei Yingjun, Hou Xinwen
The information bottleneck principle is an elegant and useful approach to representation learning. In this paper, we investigate the problem of representation learning in the conte…