activity
20172021
most citedVariational Policy Gradient Method for Reinforcement Learning with General Utilities

37 citations · 87 across the 10 of their papers we have counts for

collaborators
Showing math.OCShow all

8 papers · 1 filter

math.OC20211 cited

Zero-sum risk-sensitive continuous-time stochastic games with unbounded payoff and transition rates and Borel spaces

Junyu Zhang, Xianping Guo, Li Xia

We study a finite-horizon two-person zero-sum risk-sensitive stochastic game for continuous-time Markov chains and Borel state and action spaces, in which payoff rates, transition…

math.OC20202 cited

On the Divergence of Decentralized Non-Convex Optimization

Mingyi Hong, Siliang Zeng, Junyu Zhang +1

We study a generic class of decentralized algorithms in which agents jointly optimize the non-convex objective , while only communicating with th…

math.OC20206 cited

Generalization Bounds for Stochastic Saddle Point Problems

Junyu Zhang, Mingyi Hong, Mengdi Wang +1

This paper studies the generalization bounds for the empirical saddle point (ESP) solution to stochastic saddle point (SSP) problems. For SSP with Lipschitz continuous and strongly…

math.OC201910 cited

A Stochastic Composite Gradient Method with Incremental Variance Reduction

Junyu Zhang, Lin Xiao

We consider the problem of minimizing the composition of a smooth (nonconvex) function and a smooth vector mapping, where the inner mapping is in the form of an expectation over so…

math.OC2018

Adaptive Stochastic Variance Reduction for Subsampled Newton Method with Cubic Regularization

Junyu Zhang, Lin Xiao, Shuzhong Zhang

The cubic regularized Newton method of Nesterov and Polyak has become increasingly popular for non-convex optimization because of its capability of finding an approximate local sol…

math.OC2018

A Cubic Regularized Newton's Method over Riemannian Manifolds

Junyu Zhang, Shuzhong Zhang

In this paper we present a cubic regularized Newton's method to minimize a smooth function over a Riemannian manifold. The proposed algorithm is shown to reach a second-order -s…