1 paper
Guangzhi Sun, Xiao Zhan, Shutong Feng +2
Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety benchmarks often focus solely on the refu…