2 papers
cs.CV2026
Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior
Xiang Li, Dianbo Liu, Kenji Kawaguchi
Despite the remarkable fidelity of generative models, they frequently suffer from mode collapse. Existing strategies for enhancing diversity predominantly focus on intervening duri…
cs.CL2024
Learning diverse attacks on large language models for robust red-teaming and safety tuning
Seanie Lee, Minsu Kim, Lynn Cherif +8
Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing ef…