42 citations · 112 across the 4 of their papers we have counts for
4 papers
Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound
Lin F. Yang, Mengdi Wang
Exploration in reinforcement learning (RL) suffers from the curse of dimensionality when the state-action space is large. A common practice is to parameterize the high-dimensional…
Stochastic Primal-Dual Methods and Sample Complexity of Reinforcement Learning
Yichen Chen, Mengdi Wang
We study the online estimation of the optimal policy of a Markov decision process (MDP). We propose a class of Stochastic Primal-Dual (SPD) methods which exploit the inherent minim…
Accelerating Stochastic Composition Optimization
Mengdi Wang, Ji Liu, Ethan X. Fang
Consider the stochastic composition optimization problem where the objective is a composition of two expected-value functions. We propose a new stochastic first-order method, namel…
Stochastic Compositional Gradient Descent: Algorithms for Minimizing Compositions of Expected-Value Functions
Mengdi Wang, Ethan X. Fang, Han Liu
Classical stochastic gradient methods are well suited for minimizing expected-value objective functions. However, they do not apply to the minimization of a nonlinear function invo…