1 paper
Xiaozhe Zhang, Chaozhuo Li, Hui Liu +4
Large language models remain vulnerable to adversarial prompts that elicit harmful outputs. Existing safety paradigms typically couple red-teaming and post-training in a closed, po…