5 citations · 6 across the 6 of their papers we have counts for
1 paper · 2 filters
Seongmin Lee, Aeree Cho, Grace C. Kim +3
As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe out…