1 paper
Adel Khorramrouz, Sharon Levy
Safety guardrails in large language models(LLMs) are developed to prevent malicious users from generating toxic content at a large scale. However, these measures can inadvertently…