1 paper
John Wikman, Alexandre Proutiere, David Broman
In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that…