1 paper
Tom Bewley, Jonathan Lawry, Arthur Richards
We propose a method to capture the handling abilities of fast jet pilots in a software model via reinforcement learning (RL) from human preference feedback. We use pairwise prefere…