1 paper
Shicong Cen, Yuejie Chi
Policy gradient methods, where one searches for the policy of interest by maximizing the value functions using first-order information, become increasingly popular for sequential d…