1 paper · 1 filter
Florian Mai, David Kaczér, Nicholas Kluge Corrêa +1
Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they…