activity
20182022
most citedInstance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning

2 citations · 3 across the 3 of their papers we have counts for

collaborators

7 papers

stat.ML20221 cited

Instance-Dependent Confidence and Early Stopping for Reinforcement Learning

Koulik Khamaru, Eric Xia, Martin J. Wainwright +1

Various algorithms for reinforcement learning (RL) exhibit dramatic variation in their convergence rates as a function of problem structure. Such problem-dependent behavior is not…

stat.ML20212 cited

Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning

Koulik Khamaru, Eric Xia, Martin J. Wainwright +1

Various algorithms in reinforcement learning exhibit dramatic variability in their convergence rates and ultimate accuracy as a function of the problem structure. Such instance-spe…

stat.ML2020

Is Temporal Difference Learning Optimal? An Instance-Dependent Analysis

Koulik Khamaru, Ashwin Pananjady, Feng Ruan +2

We address the problem of policy evaluation in discounted Markov decision processes, and provide instance-dependent guarantees on the -error under a generative model.…

cs.LG2018

Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems

Dhruv Malik, Ashwin Pananjady, Kush Bhatia +3

We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-…

math.ST2018

Singularity, Misspecification, and the Convergence Rate of EM

Raaz Dwivedi, Nhat Ho, Koulik Khamaru +3

A line of recent work has analyzed the behavior of the Expectation-Maximization (EM) algorithm in the well-specified setting, in which the population likelihood is locally strongly…

stat.ML2018

Convergence guarantees for a class of non-convex and non-smooth optimization problems

Koulik Khamaru, Martin J. Wainwright

We consider the problem of finding critical points of functions that are non-convex and non-smooth. Studying a fairly broad class of such problems, we analyze the behavior of three…