activity
20182021
most citedProvably Correct Optimization and Exploration with Non-linear Policies

2 citations · 2 across the 2 of their papers we have counts for

collaborators

6 papers

cs.LG20212 cited

Provably Correct Optimization and Exploration with Non-linear Policies

Fei Feng, Wotao Yin, Alekh Agarwal +1

Policy optimization methods remain a powerful workhorse in empirical Reinforcement Learning (RL), with a focus on neural policies that can easily reason over complex and continuous…

cs.LG2020

Provably Efficient Exploration for Reinforcement Learning Using Unsupervised Learning

Fei Feng, Ruosong Wang, Wotao Yin +2

Motivated by the prevailing paradigm of using unsupervised learning for efficient exploration in reinforcement learning (RL) problems [tang2017exploration,bellemare2016unifying], w…

cs.LG2019

How Does an Approximate Model Help in Reinforcement Learning?

Fei Feng, Wotao Yin, Lin F. Yang

One of the key approaches to save samples in reinforcement learning (RL) is to use knowledge from an approximate model such as its simulator. However, how much does an approximate…

math.OC2019

Acceleration of SVRG and Katyusha X by Inexact Preconditioning

Yanli Liu, Fei Feng, Wotao Yin

Empirical risk minimization is an important class of optimization problems with many popular machine learning applications, and stochastic variance reduction methods are popular ch…

math.OC2018

AsyncQVI: Asynchronous-Parallel Q-Value Iteration for Discounted Markov Decision Processes with Near-Optimal Sample Complexity

Yibo Zeng, Fei Feng, Wotao Yin

In this paper, we propose AsyncQVI, an asynchronous-parallel Q-value iteration for discounted Markov decision processes whose transition and reward can only be sampled through a ge…

math.OC2018

A2BCD: An Asynchronous Accelerated Block Coordinate Descent Algorithm With Optimal Complexity

Robert Hannah, Fei Feng, Wotao Yin

In this paper, we propose the Asynchronous Accelerated Nonuniform Randomized Block Coordinate Descent algorithm (A2BCD), the first asynchronous Nesterov-accelerated algorithm that…