1 paper
Gerald Tesauro, Gregory R. Galperin
We present a Monte-Carlo simulation algorithm for real-time policy improvement of an adaptive controller. In the Monte-Carlo simulation, the long-term expected reward of each possi…