1 paper
Albus W. Ng, Yi Han, Jusheng Zhang +1
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for auton…