1 paper
Bruce Changlong Xu, Lan Wu
A modern post-training pipeline often writes one symbol for its policy, pi_theta, while evaluating it through two different programs: a training kernel optimized for autograd and a…