1 paper
Zachary Coalson, Beth Sohler, Aiden Gabriel +1
We identify a structural weakness in current large language model (LLM) alignment: modern refusal mechanisms are fail-open. While existing approaches encode refusal behaviors acros…