activity
20172022
most citedTheoretically Principled Trade-off between Robustness and Accuracy

921 citations · 990 across the 14 of their papers we have counts for

collaborators

23 papers

cs.LG20221 cited

Beyond the Best: Estimating Distribution Functionals in Infinite-Armed Bandits

Yifei Wang, Tavor Baharav, Yanjun Han +2

In the infinite-armed bandit problem, each arm's average reward is sampled from an unknown distribution, and each arm can be sampled further to obtain noisy estimates of the averag…

cs.LG2022

Optimal Conservative Offline RL with General Function Approximation via Augmented Lagrangian

Paria Rashidinejad, Hanlin Zhu, Kunhe Yang +2

Offline reinforcement learning (RL), which refers to decision-making from a previously-collected dataset of interactions, has received significant attention over the past years. Mu…

cs.LG20222 cited

Robust Estimation for Nonparametric Families via Generative Adversarial Networks

Banghua Zhu, Jiantao Jiao, Michael I. Jordan

We provide a general framework for designing Generative Adversarial Networks (GANs) to solve high dimensional robust statistics problems, which aim at estimating unknown parameter…

cs.LG202111 cited

MADE: Exploration via Maximizing Deviation from Explored Regions

Tianjun Zhang, Paria Rashidinejad, Jiantao Jiao +3

In online reinforcement learning (RL), efficient exploration remains particularly challenging in high-dimensional environments with sparse rewards. In low-dimensional environments,…

cs.LG20213 cited

Provably Breaking the Quadratic Error Compounding Barrier in Imitation Learning, Optimally

Nived Rajaraman, Yanjun Han, Lin F. Yang +2

We study the statistical limits of Imitation Learning (IL) in episodic Markov Decision Processes (MDPs) with a state space . We focus on the known-transition setting w…

stat.ML2021

Minimax Off-Policy Evaluation for Multi-Armed Bandits

Cong Ma, Banghua Zhu, Jiantao Jiao +1

We study the problem of off-policy evaluation in the multi-armed bandit model with bounded rewards, and develop minimax rate-optimal procedures under three settings. First, when th…