1 citations · 2 across the 14 of their papers we have counts for
1 paper · 2 filters
Ranjie Duan, Jiexi Liu, Xiaojun Jia +27
Large language models (LLMs) typically deploy safety mechanisms to prevent harmful content generation. Most current approaches focus narrowly on risks posed by malicious actors, of…