3 papers
cs.LG2024
Stable Offline Value Function Learning with Bisimulation-based Representations
Brahma S. Pavse, Yudong Chen, Qiaomin Xie +1
In reinforcement learning, offline value function learning is the procedure of using an offline dataset to estimate the expected discounted return from each state when taking actio…
cs.LG2024
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
Subhojyoti Mukherjee, Josiah P. Hanna, Qiaomin Xie +1
We study learning to learn for the multi-task structured bandit problem where the goal is to learn a near-optimal algorithm that minimizes cumulative regret. The tasks share a comm…
cs.LG2023
Multi-task Representation Learning for Pure Exploration in Bilinear Bandits
Subhojyoti Mukherjee, Qiaomin Xie, Josiah P. Hanna +1
We study multi-task representation learning for the problem of pure exploration in bilinear bandits. In bilinear bandits, an action takes the form of a pair of arms from two differ…