157 citations · 546 across the 8 of their papers we have counts for
4 papers · 1 filter
Fitted Q-Learning for Relational Domains
Srijita Das, Sriraam Natarajan, Kaushik Roy +2
We consider the problem of Approximate Dynamic Programming in relational domains. Inspired by the success of fitted Q-learning methods in propositional settings, we develop the fir…
Deep Radial-Basis Value Functions for Continuous Control
Kavosh Asadi, Neev Parikh, Ronald E. Parr +2
A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned…
Revisiting the Softmax Bellman Operator: New Benefits and New Perspective
Zhao Song, Ronald E. Parr, Lawrence Carin
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interfer…
Value Function Approximation in Noisy Environments Using Locally Smoothed Regularized Approximate Linear Programs
Gavin Taylor, Ron Parr
Recently, Petrik et al. demonstrated that L1Regularized Approximate Linear Programming (RALP) could produce value functions and policies which compared favorably to established lin…