1 paper
Michele Guida, Ruslan Shikhhamzayev, Sindhuja Penchala +4
Large language models (LLMs) can be induced to produce harmful content through multi turn strategies in which no single user message appears clearly unsafe. Existing runtime safegu…