2 papers
cs.LG2023
Value function estimation using conditional diffusion models for control
Bogdan Mazoure, Walter Talbott, Miguel Angel Bautista +3
A fairly reliable trend in deep reinforcement learning is that the performance scales with the number of parameters, provided a complimentary scaling in amount of training data. As…
cs.LG2023
Accelerating exploration and representation learning with offline pre-training
Bogdan Mazoure, Jake Bruce, Doina Precup +2
Sequential decision-making agents struggle with long horizon tasks, since solving them requires multi-step reasoning. Most reinforcement learning (RL) algorithms address this chall…