activity
20162021
most citedImproved Regret Bound and Experience Replay in Regularized Policy Iteration

2 citations · 4 across the 4 of their papers we have counts for

collaborators

11 papers

cs.LG20212 cited

Improved Regret Bound and Experience Replay in Regularized Policy Iteration

Nevena Lazic, Dong Yin, Yasin Abbasi-Yadkori +1

In this work, we study algorithms for learning in infinite-horizon undiscounted Markov decision processes (MDPs) with function approximation. We first show that the regret analysis…

cs.LG20212 cited

Optimization Issues in KL-Constrained Approximate Policy Iteration

Nevena Lazić, Botao Hao, Yasin Abbasi-Yadkori +2

Many reinforcement learning algorithms can be seen as versions of approximate policy iteration (API). While standard API often performs poorly, it has been shown that learning can…

cs.LG2020

Neural Rate Control for Video Encoding using Imitation Learning

Hongzi Mao, Chenjie Gu, Miaosen Wang +9

In modern video encoders, rate control is a critical component and has been heavily engineered. It decides how many bits to spend to encode each frame, in order to optimize the rat…

cs.LG2020

A maximum-entropy approach to off-policy evaluation in average-reward MDPs

Nevena Lazic, Dong Yin, Mehrdad Farajtabar +4

This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs). For MDPs that are ergodic and linear…

cs.LG2020

Robotic Table Tennis with Model-Free Reinforcement Learning

Wenbo Gao, Laura Graesser, Krzysztof Choromanski +5

We propose a model-free algorithm for learning efficient policies capable of returning table tennis balls by controlling robot joints at a rate of 100Hz. We demonstrate that evolut…

cs.LG2020

Adaptive Approximate Policy Iteration

Botao Hao, Nevena Lazic, Yasin Abbasi-Yadkori +2

Model-free reinforcement learning algorithms combined with value function approximation have recently achieved impressive performance in a variety of application domains. However,…