1 paper
Mohammad Omar Khursheed, Baram Sosis, Fabien Roger
Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of AI R&D where refusals would…