activity
20182026
most citedVariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning

65 citations · 124 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2024

Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models

Zhejun Zhang, Peter Karkus, Maximilian Igl +4

Traffic simulation aims to learn a policy for traffic agents that, when unrolled in closed-loop, faithfully recovers the joint distribution of trajectories observed in the real wor…

cs.LG2023

Hierarchical Imitation Learning for Stochastic Environments

Maximilian Igl, Punit Shah, Paul Mougin +5

Many applications of imitation learning require the agent to generate the full distribution of behaviour observed in the training data. For example, to evaluate the safety of auton…

cs.LG2022

Symphony: Learning Realistic and Diverse Agents for Autonomous Driving Simulation

Maximilian Igl, Daewoo Kim, Alex Kuefler +7

Simulation is a crucial tool for accelerating the development of autonomous vehicles. Making simulation realistic requires models of the human road users who interact with such car…

cs.LG2020

My Body is a Cage: the Role of Morphology in Graph-Based Incompatible Control

Vitaly Kurin, Maximilian Igl, Tim Rocktäschel +2

Multitask Reinforcement Learning is a promising way to obtain models with better performance, generalisation, data efficiency, and robustness. Most existing work is limited to comp…

cs.LG201958 cited

Generalization in Reinforcement Learning with Selective Noise Injection and Information Bottleneck

Maximilian Igl, Kamil Ciosek, Yingzhen Li +4

The ability for policies to generalize to new environments is key to the broad application of RL agents. A promising approach to prevent an agent's policy from overfitting to a lim…

cs.LG201965 cited

VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning

Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl +4

Trading off exploration and exploitation in an unknown environment is key to maximising expected return during learning. A Bayes-optimal policy, which does so optimally, conditions…