1 paper
Giorgio Maria Cavallazzi, Miguel Pérez-Cuadrado, Miguel Pérez-Cuadrado +1
Reinforcement-learning controllers optimise specified rewards, but in physical systems those rewards often capture only part of the true control objective. Three mechanisms through…