5 citations · 5 across the 3 of their papers we have counts for
1 paper · 1 filter
Tom Bewley, Jonathan Lawry, Arthur Richards
We propose a method to capture the handling abilities of fast jet pilots in a software model via reinforcement learning (RL) from human preference feedback. We use pairwise prefere…