7 citations · 7 across the 1 of their papers we have counts for
1 paper
Stephen Casper, Jason Lin, Joe Kwon +2
Deploying large language models (LMs) can pose hazards from harmful outputs such as toxic or false text. Prior work has introduced automated tools that elicit harmful outputs to id…