2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Arnaud Lequen, Clément Legrand-Lixon, Léo Saulières
We present Pareto Q-Learning with Reward Machines (PQLRM), a multi-objective reinforcement learning algorithm for tasks whose reward structure is specified by a set of reward machi…