1 paper · 1 filter
Leonard Dung, Florian Mai
AI alignment research aims to develop techniques to ensure that AI systems do not cause harm. However, every alignment technique has failure modes, which are conditions in which th…