activity
20162022
most citedVariational Policy Gradient Method for Reinforcement Learning with General Utilities

37 citations · 56 across the 13 of their papers we have counts for

collaborators
Showing math.OCShow all

5 papers · 1 filter

math.OC2021

Convergence Rates of Average-Reward Multi-agent Reinforcement Learning via Randomized Linear Programming

Alec Koppel, Amrit Singh Bedi, Bhargav Ganguly +1

In tabular multi-agent reinforcement learning with average-cost criterion, a team of agents sequentially interacts with the environment and observes local incentives. We focus on t…

math.OC2019

Nonstationary Nonparametric Online Learning: Balancing Dynamic Regret and Model Parsimony

Amrit Singh Bedi, Alec Koppel, Ketan Rajawat +1

An open challenge in supervised learning is \emph{conceptual drift}: a data point begins as classified according to one label, but over time the notion of that label changes. Beyon…

math.OC2019

Global Convergence of Policy Gradient Methods to (Almost) Locally Optimal Policies

Kaiqing Zhang, Alec Koppel, Hao Zhu +1

Policy gradient (PG) methods are a widely used reinforcement learning methodology in many applications such as video games, autonomous driving, and robotics. In spite of its empiri…

math.OC2019

Nonparametric Compositional Stochastic Optimization for Risk-Sensitive Kernel Learning

Amrit Singh Bedi, Alec Koppel, Ketan Rajawat +1

In this work, we address optimization problems where the objective function is a nonlinear function of an expected value, i.e., compositional stochastic {strongly convex programs}.…

math.OC20172 cited

Asynchronous Decentralized Stochastic Optimization in Heterogeneous Networks

Amrit Singh Bedi, Alec Koppel, Ketan Rajawat

We consider expected risk minimization in multi-agent systems comprised of distinct subsets of agents operating without a common time-scale. Each individual in the network is charg…