1 paper · 1 filter
Sathwik Karnik, Somil Bansal
Large language models (LLMs) are now ubiquitous in everyday tools, raising urgent safety concerns about their tendency to generate harmful content. The dominant safety approach --…