1 paper
Anas Barakat, Ilyas Fatkhullin, Niao He
We consider the reinforcement learning (RL) problem with general utilities which consists in maximizing a function of the state-action occupancy measure. Beyond the standard cumula…