1 paper
Simon Dima, Simon Fischer, Jobst Heitzig +1
In dynamic programming and reinforcement learning, the policy for the sequential decision making of an agent in a stochastic environment is usually determined by expressing the goa…