1 paper · 1 filter
Seanie Lee, Minsu Kim, Lynn Cherif +8
Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing ef…