1 paper
Tomonori Sadamoto, Takashi Tanaka
We study the policy gradient method (PGM) for the linear quadratic Gaussian (LQG) dynamic output-feedback control problem using an input-output-history (IOH) representation of the…