1 paper
Hippolyte Bourel, Anders Jonsson, Odalric-Ambrym Maillard +2
We study reinforcement learning (RL) for decision processes with non-Markovian reward, in which high-level knowledge of the task in the form of reward machines is available to the…