activity
20172026
most citedMulti-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces

43 citations · 103 across the 59 of their papers we have counts for

collaborators
Showing cs.AIShow all

13 papers · 1 filter

cs.AI2026

From Noise to Control: Parameterized Diffusion Policies

Renhao Zhang, Haotian Fu, Mingxi Jia +3

We propose Parameterized Diffusion Policy (PDP), a framework for learning diffusion policies conditioned on low-dimensional, continuous parameters embedded in a learned behavior ma…

cs.AI2025

Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning

Martin Klissarov, Akhil Bagaria, Ziyan Luo +3

Developing agents capable of exploring, planning and learning in complex open-ended environments is a grand challenge in artificial intelligence (AI). Hierarchical reinforcement le…

cs.AI2025

Automating Curriculum Learning for Reinforcement Learning using a Skill-Based Bayesian Network

Vincent Hsiao, Mark Roberts, Laura M. Hiatt +2

A major challenge for reinforcement learning is automatically generating curricula to reduce training time or improve performance in some target task. We introduce SEBNs (Skill-Env…

cs.AI2024

Latent-Predictive Empowerment: Measuring Empowerment without a Simulator

Andrew Levy, Alessandro Allievi, George Konidaris

Empowerment has the potential to help agents learn large skillsets, but is not yet a scalable solution for training general-purpose agents. Recent empowerment methods learn diverse…

cs.AI2023

MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds

William Hill, Ireton Liu, Anita De Mello Koch +4

We propose a new benchmark for planning tasks based on the Minecraft game. Our benchmark contains 45 tasks overall, but also provides support for creating both propositional and nu…

cs.AI2023

Exploiting Contextual Structure to Generate Useful Auxiliary Tasks

Benedict Quartey, Ankit Shah, George Konidaris

Reinforcement learning requires interaction with an environment, which is expensive for robots. This constraint necessitates approaches that work with limited environmental interac…