1 paper · 1 filter
Perry Dong, Suvir Mirchandani, Dorsa Sadigh +1
The ability to learn from large batches of autonomously collected data for policy improvement -- a paradigm we refer to as batch online reinforcement learning -- holds the promise…