1 paper · 1 filter
Yuxiao Lu, Arunesh Sinha, Pradeep Varakantham
Large Language Models (LLMs) generating unsafe responses to toxic prompts is a significant issue in their applications. While various efforts aim to address this safety concern, pr…