1 paper · 1 filter
Luca Marzari, Priya L. Donti, Changliu Liu +1
We present ε-retrain, an exploration strategy encouraging a behavioral preference while optimizing policies with monotonic improvement guarantees. To this end, we intro…