activity
20172024
most citedModel-Based Reinforcement Learning with Value-Targeted Regression

69 citations · 393 across the 21 of their papers we have counts for

collaborators
Showing math.OCShow all

7 papers · 1 filter

math.OC20206 cited

Generalization Bounds for Stochastic Saddle Point Problems

Junyu Zhang, Mingyi Hong, Mengdi Wang +1

This paper studies the generalization bounds for the empirical saddle point (ESP) solution to stochastic saddle point (SSP) problems. For SSP with Lipschitz continuous and strongly…

math.OC2018

A Single Time-Scale Stochastic Approximation Method for Nested Stochastic Optimization

Saeed Ghadimi, Andrzej Ruszczyński, Mengdi Wang

We study constrained nested stochastic optimization problems in which the objective function is a composition of two smooth functions whose exact values and derivatives are not ava…

math.OC2018

Adaptive Low-Nonnegative-Rank Approximation for State Aggregation of Markov Chains

Yaqi Duan, Mengdi Wang, Zaiwen Wen +1

This paper develops a low-nonnegative-rank approximation method to identify the state aggregation structure of a finite-state Markov chain under an assumption that the state space…

math.OC2018

Near-Optimal Time and Sample Complexities for Solving Discounted Markov Decision Process with a Generative Model

Aaron Sidford, Mengdi Wang, Xian Wu +2

In this paper we consider the problem of computing an -optimal policy of a discounted Markov Decision Process (DMDP) provided we can only access its transition function through…

math.OC2018

Improved Sample Complexity for Stochastic Compositional Variance Reduced Gradient

Tianyi Lin, Chenyou Fan, Mengdi Wang +1

Convex composition optimization is an emerging topic that covers a wide range of applications arising from stochastic optimal control, reinforcement learning and multi-stage stocha…

math.OC2018

Approximation Methods for Bilevel Programming

Saeed Ghadimi, Mengdi Wang

In this paper, we study a class of bilevel programming problem where the inner objective function is strongly convex. More specifically, under some mile assumptions on the partial…