1 paper
Lily Zhang, Xianling Zhang
Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: ca…