activity
20182022
most citedSample Efficient Reinforcement Learning In Continuous State Spaces: A Perspective Beyond Linearity

3 citations · 3 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG2022

How Does Adaptive Optimization Impact Local Neural Network Geometry?

Kaiqi Jiang, Dhruv Malik, Yuanzhi Li

Adaptive optimization methods are well known to achieve superior convergence relative to vanilla gradient methods. The traditional viewpoint in optimization, particularly in convex…

stat.ML2022

Complete Policy Regret Bounds for Tallying Bandits

Dhruv Malik, Yuanzhi Li, Aarti Singh

Policy regret is a well established notion of measuring the performance of an online learning algorithm against an adaptive adversary. We study restrictions on the adversary that e…

cs.LG20213 cited

Sample Efficient Reinforcement Learning In Continuous State Spaces: A Perspective Beyond Linearity

Dhruv Malik, Aldo Pacchiano, Vishwak Srinivasan +1

Reinforcement learning (RL) is empirically successful in complex nonlinear Markov decision processes (MDPs) with continuous state spaces. By contrast, the majority of theoretical R…

cs.LG2021

When Is Generalizable Reinforcement Learning Tractable?

Dhruv Malik, Yuanzhi Li, Pradeep Ravikumar

Agents trained by reinforcement learning (RL) often fail to generalize beyond the environment they were trained in, even when presented with new scenarios that seem similar to the…

cs.LG2018

Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems

Dhruv Malik, Ashwin Pananjady, Kush Bhatia +3

We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-…

cs.AI2018

An Efficient, Generalized Bellman Update For Cooperative Inverse Reinforcement Learning

Dhruv Malik, Malayandi Palaniappan, Jaime F. Fisac +3

Our goal is for AI systems to correctly identify and act according to their human user's objectives. Cooperative Inverse Reinforcement Learning (CIRL) formalizes this value alignme…