921 citations · 990 across the 14 of their papers we have counts for
23 papers
Beyond the Best: Estimating Distribution Functionals in Infinite-Armed Bandits
Yifei Wang, Tavor Baharav, Yanjun Han +2
In the infinite-armed bandit problem, each arm's average reward is sampled from an unknown distribution, and each arm can be sampled further to obtain noisy estimates of the averag…
Optimal Conservative Offline RL with General Function Approximation via Augmented Lagrangian
Paria Rashidinejad, Hanlin Zhu, Kunhe Yang +2
Offline reinforcement learning (RL), which refers to decision-making from a previously-collected dataset of interactions, has received significant attention over the past years. Mu…
Robust Estimation for Nonparametric Families via Generative Adversarial Networks
Banghua Zhu, Jiantao Jiao, Michael I. Jordan
We provide a general framework for designing Generative Adversarial Networks (GANs) to solve high dimensional robust statistics problems, which aim at estimating unknown parameter…
MADE: Exploration via Maximizing Deviation from Explored Regions
Tianjun Zhang, Paria Rashidinejad, Jiantao Jiao +3
In online reinforcement learning (RL), efficient exploration remains particularly challenging in high-dimensional environments with sparse rewards. In low-dimensional environments,…
Provably Breaking the Quadratic Error Compounding Barrier in Imitation Learning, Optimally
Nived Rajaraman, Yanjun Han, Lin F. Yang +2
We study the statistical limits of Imitation Learning (IL) in episodic Markov Decision Processes (MDPs) with a state space . We focus on the known-transition setting w…
Minimax Off-Policy Evaluation for Multi-Armed Bandits
Cong Ma, Banghua Zhu, Jiantao Jiao +1
We study the problem of off-policy evaluation in the multi-armed bandit model with bounded rewards, and develop minimax rate-optimal procedures under three settings. First, when th…