activity
20172024
most citedNeural Contextual Bandits with Deep Representation and Shallow Exploration

18 citations · 63 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

20 papers · 1 filter

cs.LG2024

Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation

Yihong Guo, Yixuan Wang, Yuanyuan Shi +2

Training a policy in a source domain for deployment in the target domain under a dynamics shift can be challenging, often resulting in performance degradation. Previous work tackle…

cs.LG2024

Robust Offline Reinforcement Learning with Linearly Structured f-Divergence Regularization

Cheng Tang, Zhishuai Liu, Pan Xu

The Robust Regularized Markov Decision Process (RRMDP) is proposed to learn policies robust to dynamics shifts by adding regularization to the transition dynamics in the value func…

cs.LG20241 cited

More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling

Haque Ishfaq, Yixin Tan, Yu Yang +5

Thompson sampling (TS) is one of the most popular exploration techniques in reinforcement learning (RL). However, most TS algorithms with theoretical guarantees are difficult to im…

cs.LG2024

Optimal Batched Linear Bandits

Xuanfei Ren, Tianyuan Jin, Pan Xu

We introduce the E algorithm for the batched linear bandit problem, incorporating an Explore-Estimate-Eliminate-Exploit framework. With a proper choice of exploration rate, we…

cs.LG2024

Randomized Exploration in Cooperative Multi-Agent Reinforcement Learning

Hao-Lun Hsu, Weixin Wang, Miroslav Pajic +1

We present the first study on provably efficient randomized exploration in cooperative multi-agent reinforcement learning (MARL). We propose a unified algorithm framework for rando…

cs.LG2024

Minimax Optimal and Computationally Efficient Algorithms for Distributionally Robust Offline Reinforcement Learning

Zhishuai Liu, Pan Xu

Distributionally robust offline reinforcement learning (RL), which seeks robust policy training against environment perturbation by modeling dynamics uncertainty, calls for functio…