1 paper
Soyeon Park, Seogyeong Jeong, Sunwoo Kim +1
Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. T…