57 citations · 230 across the 22 of their papers we have counts for
3 papers · 1 filter
Understanding the Pathologies of Approximate Policy Evaluation when Combined with Greedification in Reinforcement Learning
Kenny Young, Richard S. Sutton
Despite empirical success, the theory of reinforcement learning (RL) with value function approximation remains fundamentally incomplete. Prior work has identified a variety of path…
Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
Katya Kudashkina, Patrick M. Pilarski, Richard S. Sutton
Intelligent assistants that follow commands or answer simple questions, such as Siri and Google search, are among the most economically important applications of AI. Future convers…
Inverse Policy Evaluation for Value-based Sequential Decision-making
Alan Chan, Kris de Asis, Richard S. Sutton
Value-based methods for reinforcement learning lack generally applicable ways to derive behavior from a value function. Many approaches involve approximate value iteration (e.g., $…