1 paper
Zachary Coalson, Jeonghyun Woo, Chris S. Lin +8
We study a new vulnerability in commercial-scale safety-aligned large language models (LLMs): their refusal to generate harmful responses can be broken by flipping only a few bits…