Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
David Dobre, Mehrnaz Mofakhami, Sophie Xhonneux +2
Many safety post-training methods for large language models (LLMs) are designed to modify the model's behaviour from producing unsafe answers to issuing refusals. However, such dis…
cs.CL2025
Learning diverse attacks on large language models for robust red-teaming and safety tuning
Seanie Lee, Minsu Kim, Lynn Cherif +8
Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing ef…