1 paper
Jingnan Zheng, Han Wang, An Zhang +3
Large Language Models (LLMs) can elicit unintended and even harmful content when misaligned with human values, posing severe risks to users and society. To mitigate these risks, cu…