4 papers
Reinforcement Learning From State and Temporal Differences
Lex Weaver, Jonathan Baxter
TD() with function approximation has proved empirically successful for some complex reinforcement learning problems. For linear approximation, TD() has been shown to minimi…
A result relating convex n-widths to covering numbers with some applications to neural networks
Jonathan Baxter, Peter Bartlett
In general, approximating classes of functions defined over high-dimensional input spaces by linear combinations of a fixed set of basis functions or ``features'' is known to be ha…
Scaling Internal-State Policy-Gradient Methods for POMDPs
Douglas Aberdeen, Jonathan Baxter
Policy-gradient methods have received increased attention recently as a mechanism for learning to act in partially observable environments. They have shown promise for problems adm…
Reinforcement Learning in POMDP's via Direct Gradient Ascent
Jonathan Baxter, Peter L. Bartlett
This paper discusses theoretical and experimental aspects of gradient-based approaches to the direct optimization of policy performance in controlled POMDPs. We introduce GPOMDP, a…