1 paper · 1 filter
Seongmin Lee, Aeree Cho, Grace C. Kim +3
As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe out…