2 citations · 4 across the 4 of their papers we have counts for
11 papers
Improved Regret Bound and Experience Replay in Regularized Policy Iteration
Nevena Lazic, Dong Yin, Yasin Abbasi-Yadkori +1
In this work, we study algorithms for learning in infinite-horizon undiscounted Markov decision processes (MDPs) with function approximation. We first show that the regret analysis…
Optimization Issues in KL-Constrained Approximate Policy Iteration
Nevena Lazić, Botao Hao, Yasin Abbasi-Yadkori +2
Many reinforcement learning algorithms can be seen as versions of approximate policy iteration (API). While standard API often performs poorly, it has been shown that learning can…
Neural Rate Control for Video Encoding using Imitation Learning
Hongzi Mao, Chenjie Gu, Miaosen Wang +9
In modern video encoders, rate control is a critical component and has been heavily engineered. It decides how many bits to spend to encode each frame, in order to optimize the rat…
A maximum-entropy approach to off-policy evaluation in average-reward MDPs
Nevena Lazic, Dong Yin, Mehrdad Farajtabar +4
This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs). For MDPs that are ergodic and linear…
Robotic Table Tennis with Model-Free Reinforcement Learning
Wenbo Gao, Laura Graesser, Krzysztof Choromanski +5
We propose a model-free algorithm for learning efficient policies capable of returning table tennis balls by controlling robot joints at a rate of 100Hz. We demonstrate that evolut…
Adaptive Approximate Policy Iteration
Botao Hao, Nevena Lazic, Yasin Abbasi-Yadkori +2
Model-free reinforcement learning algorithms combined with value function approximation have recently achieved impressive performance in a variety of application domains. However,…