27 citations · 46 across the 5 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2022★ 1 cited
On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs
Yi Wan, Richard S. Sutton
We show two average-reward off-policy control algorithms, Differential Q-learning (Wan, Naik, & Sutton 2021a) and RVI Q-learning (Abounadi Bertsekas & Borkar 2001), converge in wea…
cs.LG2022★ 1 cited
Doubly-Asynchronous Value Iteration: Making Value Iteration Asynchronous in Actions
Tian Tian, Kenny Young, Richard S. Sutton
Value iteration (VI) is a foundational dynamic programming method, important for learning and planning in optimal control and reinforcement learning. VI proceeds in batches, where…
cs.LG2021
Learning Agent State Online with Recurrent Generate-and-Test
Amir Samani, Richard S. Sutton
Learning continually and online from a continuous stream of data is challenging, especially for a reinforcement learning agent with sequential data. When the environment only provi…