1 paper · 1 filter
Siddharth Sai, Xiaofei Wen, Muhao Chen
Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrails typically rely on single-…