4 papers
On the Convergence of Moral Self-Correction in Large Language Models
Guangliang Liu, Haitao Mao, Bochuan Cao +4
Large Language Models (LLMs) are able to improve their responses when instructed to do so, a capability known as self-correction. When instructions provide only a general and abstr…
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
Guangliang Liu, Bocheng Chen, Han Zi +2
Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender…
Discourse Heuristics For Paradoxically Moral Self-Correction
Guangliang Liu, Zimo Qi, Xitong Zhang +1
Moral self-correction has emerged as a promising approach for aligning the output of Large Language Models (LLMs) with human moral values. However, moral self-correction techniques…
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
Zhiyu Xue, Guangliang Liu, Bocheng Chen +2
The security of Large Language Models (LLMs) has become an important research topic since the emergence of ChatGPT. Though there have been various effective methods to defend again…