1 paper
Agam Goyal, Vedant Rathi, William Yeh +3
Large language models (LLMs) are now ubiquitous in user-facing applications, yet they still generate undesirable toxic outputs, including profanity, vulgarity, and derogatory remar…