1 paper
Wouter Jongeneel, Daniel Kuhn, Mengmeng Li
Motivated by policy gradient methods in the context of reinforcement learning, we identify a large deviation rate function for the iterates generated by stochastic gradient descent…