8 papers
How Far Can Sharpness and Complexity Jointly Explain Generalization?
Ziyu Cheng, Xitong Zhang, Longxiu Huang +1
Sharpness and complexity are two central factors in the generalization analysis of deep neural networks. Existing quantitative evaluations of generalization measures have largely f…
From Training to Generalization: Improving Moral Reasoning Through Pragmatic Inference
Guangliang Liu, Xi Chen, Bocheng Chen +3
Although moral reasoning has emerged as a promising research direction for large language models (LLMs), a persistent generalization challenge remains: LLMs often achieve strong pe…
Learning to Diagnose and Correct Moral Errors: Beyond Shallow Heuristics in Moral Alignment
Bocheng Chen, Xi Chen, Han Zi +5
Existing approaches to moral value alignment are primarily set out to align LLMs' generation with the distributions of morally appropriate language, which has seen good progress. H…
Self-correction is Not An Innate Capability in Language Models
Guangliang Liu, Zimo Qi, Xitong Zhang +2
Although there has been growing interest in the self-correction capability of Large Language Models (LLMs), there are varying conclusions about its effectiveness. Prior research ha…
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
Guangliang Liu, Bocheng Chen, Han Zi +2
Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender…
Discourse Heuristics For Paradoxically Moral Self-Correction
Guangliang Liu, Zimo Qi, Xitong Zhang +1
Moral self-correction has emerged as a promising approach for aligning the output of Large Language Models (LLMs) with human moral values. However, moral self-correction techniques…