1 paper
Douglas Aberdeen, Jonathan Baxter
Policy-gradient methods have received increased attention recently as a mechanism for learning to act in partially observable environments. They have shown promise for problems adm…