Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023
A Strong Baseline for Batch Imitation Learning
Matthew Smith, Lucas Maystre, Zhenwen Dai +1
Imitation of expert behaviour is a highly desirable and safe approach to the problem of sequential decision making. We provide an easy-to-implement, novel algorithm for imitation l…
cs.LG2023
Why Target Networks Stabilise Temporal Difference Methods
Mattie Fellows, Matthew J. A. Smith, Shimon Whiteson
Integral to recent successes in deep reinforcement learning has been a class of temporal difference methods that use infrequently updated target values for policy evaluation in a M…