17 citations · 17 across the 2 of their papers we have counts for
1 paper · 1 filter
Carlo Alfano, Sebastian Towers, Silvia Sapora +2
Policy Mirror Descent (PMD) is a popular framework in reinforcement learning, serving as a unifying perspective that encompasses numerous algorithms. These algorithms are derived t…