1 paper
Neal Mangaokar, Ashish Hooda, Jihye Choi +4
Large language models (LLMs) are typically aligned to be harmless to humans. Unfortunately, recent work has shown that such models are susceptible to automated jailbreak attacks th…