1 paper
Vivek Farias, Joren Gijsbrechts, Aryan Khojandi +2
Simulating a single trajectory of a dynamical system under some state-dependent policy is a core bottleneck in policy optimization (PO) algorithms. The many inherently serial polic…