1 citations · 1 across the 1 of their papers we have counts for
1 paper
Zhihao Lin
Gaussian policies have dominated continuous control in deep reinforcement learning (RL), yet they suffer from a fundamental mismatch: their unbounded support requires ad-hoc squash…