1 paper
Qiuhua Pan, Yukai Shen, Liwei Zhang +2
We study the sample complexity of policy gradient for log-growth control -- the problem of learning, from observed state transitions, a feedback gain that optimally stabilizes a sc…