1 citations · 1 across the 4 of their papers we have counts for
4 papers
On the Convergence of Bounded Agents
David Abel, André Barreto, Hado van Hasselt +3
When has an agent converged? Standard models of the reinforcement learning problem give rise to a straightforward definition of convergence: An agent converges when its behavior or…
Exploration via Epistemic Value Estimation
Simon Schmitt, John Shawe-Taylor, Hado van Hasselt
How to efficiently explore in reinforcement learning is an open problem. Many exploration algorithms employ the epistemic uncertainty of their own value predictions -- for instance…
Learning How to Infer Partial MDPs for In-Context Adaptation and Exploration
Chentian Jiang, Nan Rosemary Ke, Hado van Hasselt
To generalize across tasks, an agent should acquire knowledge from past tasks that facilitate adaptation and exploration in future tasks. We focus on the problem of in-context adap…
Optimistic Meta-Gradients
Sebastian Flennerhag, Tom Zahavy, Brendan O'Donoghue +3
We study the connection between gradient-based meta-learning and convex op-timisation. We observe that gradient descent with momentum is a special case of meta-gradients, and build…