24 citations · 24 across the 2 of their papers we have counts for
1 paper · 1 filter
Alwin Peng, Julian Michael, Henry Sleight +2
As large language models (LLMs) grow more powerful, ensuring their safety against misuse becomes crucial. While researchers have focused on developing robust defenses, no method ha…