25 citations · 57 across the 4 of their papers we have counts for
1 paper · 2 filters
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner +9
Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior - for example, to evaluate wh…