activity
20212023
most citedHybrid RL: Using Both Offline and Online Data Can Make RL Efficient

7 citations · 12 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2023

Contextual Bandits and Imitation Learning via Preference-Based Active Queries

Ayush Sekhari, Karthik Sridharan, Wen Sun +1

We consider the problem of contextual bandits and imitation learning, where the learner lacks direct knowledge of the executed action's reward. Instead, the learner can actively qu…

cs.LG2022★ 7 cited

Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Yuda Song, Yifei Zhou, Ayush Sekhari +3

We consider a hybrid reinforcement learning setting (Hybrid RL), in which an agent has access to an offline dataset and the ability to collect experience via real-world online inte…

cs.LG2022★ 1 cited

PAC Reinforcement Learning for Predictive State Representations

Wenhao Zhan, Masatoshi Uehara, Wen Sun +1

In this paper we study online Reinforcement Learning (RL) in partially observable dynamical systems. We focus on the Predictive State Representations (PSRs) model, which is an expr…

cs.LG2022★ 1 cited

Learning Bellman Complete Representations for Offline Policy Evaluation

Jonathan D. Chang, Kaiwen Wang, Nathan Kallus +1

We study representation learning for Offline Reinforcement Learning (RL), focusing on the important task of Offline Policy Evaluation (OPE). Recent work shows that, in contrast to…

cs.LG2022

Computationally Efficient PAC RL in POMDPs with Latent Determinism and Conditional Embeddings

Masatoshi Uehara, Ayush Sekhari, Jason D. Lee +2

We study reinforcement learning with function approximation for large-scale Partially Observable Markov Decision Processes (POMDPs) where the state space and observation space are…

cs.LG2022★ 3 cited

Provably Efficient Reinforcement Learning in Partially Observable Dynamical Systems

Masatoshi Uehara, Ayush Sekhari, Jason D. Lee +2

We study Reinforcement Learning for partially observable dynamical systems using function approximation. We propose a new \textit{Partially Observable Bilinear Actor-Critic framewo…